wesg comments on SAE reconstruction errors are (empirically) pathological

wesg 29 Mar 2024 21:30 UTC
9 points
0
This is a great comment! The basic argument makes sense to me, though based on how much variability there is in this plot, I think the story is more complicated. Specifically, I think your theory predicts that the SAE reconstructed KL should always be out on the tail, and these random perturbations should have low variance in their effect on KL.
I will do some follow up experiments to test different versions of this story.