I definitely think so. The Turing test is a very hard target to hit, and we don’t really have a good idea how to measure IQ, knowledge, human-likeness, etc. I notice a lot of confusion, anthropomorphizing, bad analogies, etc in public discourse right now. To me it feels like the conversation is at a level where we need more precise measures that are human and machine compatible. Benchmarks based on specific tasks (as found in AI papers) don’t cut it.
(ep status: speculative)
Potentially, AI safety folks are better positioned to work on these foundational issues than traditional academics, who are very focused on capabilities and applications right now.
I definitely think so. The Turing test is a very hard target to hit, and we don’t really have a good idea how to measure IQ, knowledge, human-likeness, etc. I notice a lot of confusion, anthropomorphizing, bad analogies, etc in public discourse right now. To me it feels like the conversation is at a level where we need more precise measures that are human and machine compatible. Benchmarks based on specific tasks (as found in AI papers) don’t cut it.
(ep status: speculative) Potentially, AI safety folks are better positioned to work on these foundational issues than traditional academics, who are very focused on capabilities and applications right now.