Diego Caples comments on Scaling Sparse Feature Circuit Finding to Gemma 9B

Diego Caples 13 Jan 2025 5:39 UTC
4 points
0
Thanks! We use mean ablation because this lets us create circuits including only the things which change between examples in a task. So, for example, in the code task, our circuits do not need “is python” latents, as these latents are consistent across all samples. Were we to zero ablate, we would need every single SAE latent necessary for every part of the task. This includes things which were consistent across all tasks! This means many latents which we don’t really care about are included in our circuits.