A fire alarm creates common knowledge, in the you-know-I-know sense, that there is a fire; after which it is socially safe to react. When the fire alarm goes off, you know that everyone else knows there is a fire, you know you won’t lose face if you proceed to exit the building.
If I have a predetermined set of tests, this could serve as a fire alarm, but only if you’ve successfully built a consensus that it is one. This is hard, and the consensus would need to be quite strong. To avoid ambiguity, the test itself would need to be demonstrably resistant to being clever Hans’ed. Otherwise it would be just another milestone.
In There’s No Fire Alarm for Artificial General Intelligence Eliezer argues:
If I have a predetermined set of tests, this could serve as a fire alarm, but only if you’ve successfully built a consensus that it is one. This is hard, and the consensus would need to be quite strong. To avoid ambiguity, the test itself would need to be demonstrably resistant to being clever Hans’ed. Otherwise it would be just another milestone.
I very much agree.