Curated! I think it’s generally great when people explain what they’re doing and why in way legibile to those not working on it. Great because it let’s others potentially get involved, build on it, expose flaws or omissions, etc. This one seems particularly clear and well written. While I haven’t read all of the research, nor am I particularly qualified to comment on it, I like the idea of a principled/systematic approach behind, in comparison to a lot of work that isn’t coming on a deeper, bigger, framework.
(While I’m here though, I’ll add a link to Dmitry Vaintrob’s comment that Jacob Hilton described as “best critique of ARC’s research agenda that I have read since we started working on heuristic explanations”. Eliciting such feedback is the kind of good thing that comes out of up writing agendas – it’s possible or likely Dmitry was already tracking the work and already had these critiques, but a post like this seems like a good way to propagate them and have a public back and forth.)
Roughly speaking, if the scalability of an algorithm depends on unknown empirical contingencies (such as how advanced AI systems generalize), then we try to make worst-case assumptions instead of attempting to extrapolate from today’s systems.
I like this attitude. The human standard, I think often in alignment work too, is to argue why one’s plan will work and find stories for that, and adopting the methodology of the opposite, especially given the unknowns, is much needed in alignment work.
Overall, this is neat. Kudos to Jacob (and rest of the team) for taking the time to put this all together. Doesn’t seem all that quick to write, and I think it’d be easy to think they ought to not take time out off from further object-level research to write it. Thanks!
Curated! I think it’s generally great when people explain what they’re doing and why in way legibile to those not working on it. Great because it let’s others potentially get involved, build on it, expose flaws or omissions, etc. This one seems particularly clear and well written. While I haven’t read all of the research, nor am I particularly qualified to comment on it, I like the idea of a principled/systematic approach behind, in comparison to a lot of work that isn’t coming on a deeper, bigger, framework.
(While I’m here though, I’ll add a link to Dmitry Vaintrob’s comment that Jacob Hilton described as “best critique of ARC’s research agenda that I have read since we started working on heuristic explanations”. Eliciting such feedback is the kind of good thing that comes out of up writing agendas – it’s possible or likely Dmitry was already tracking the work and already had these critiques, but a post like this seems like a good way to propagate them and have a public back and forth.)
I like this attitude. The human standard, I think often in alignment work too, is to argue why one’s plan will work and find stories for that, and adopting the methodology of the opposite, especially given the unknowns, is much needed in alignment work.
Overall, this is neat. Kudos to Jacob (and rest of the team) for taking the time to put this all together. Doesn’t seem all that quick to write, and I think it’d be easy to think they ought to not take time out off from further object-level research to write it. Thanks!