Most end-to-end “alignment plans” are bad because research will be incremental. For example, Superalignment’s impact will mostly come from adapting to the next ~3 years of AI discoveries and working on relevant subproblems like interp, rather than creating a superhuman alignment researcher.
Most end-to-end “alignment plans” are bad because research will be incremental. For example, Superalignment’s impact will mostly come from adapting to the next ~3 years of AI discoveries and working on relevant subproblems like interp, rather than creating a superhuman alignment researcher.