An important thing to emphasize with control arguments is that it seems quite unlikely that control arguments can be made workable for very superhuman models. (At least for the notion of “control arguments” which can be readily assessed with non-insane capability evaluations.)
An important thing to emphasize with control arguments is that it seems quite unlikely that control arguments can be made workable for very superhuman models. (At least for the notion of “control arguments” which can be readily assessed with non-insane capability evaluations.)