Neel Nanda comments on dirk’s Shortform

Neel Nanda 4 Sep 2024 13:17 UTC
4 points
2
These are LLM generated labels, there are no “real” labels (because they’re expensive!). Especially in our demo, Neuronpedia made them with gpt 3.5 which is kinda dumb.

I mostly think they’re much better than nothing, but shouldn’t be trusted, and I’m glad our demo makes this apparent to people! I’m excited about work to improve autointerp, though unfortunately the easiest way is to use a better model, which gets expensive