The most similar analysis tool I’m aware of is called an activation atlas (https://distill.pub/2019/activation-atlas/), though I’ve only seen it applied to visual networks. Would love to see it used on language models!
The most similar analysis tool I’m aware of is called an activation atlas (https://distill.pub/2019/activation-atlas/), though I’ve only seen it applied to visual networks. Would love to see it used on language models!