ryan_greenblatt comments on RSPs are pauses done right

ryan_greenblatt 18 Oct 2023 4:17 UTC
LW: 25 AF: 18
22
AF
I happen to think that the Anthropic RSP is fine for what it is, but it just doesn’t actually make any interesting claims yet. The key thing is that they’re committing to actually having an ASL-4 criteria and safety argument in the future. From my perspective, the Anthropic RSP effectively is an outline for the sort of thing an RSP could be (run evals, have safety buffer, assume continuity, etc) as well as a commitment to finish the key parts of the RSP later. This seems ok to me.

I would preferred if they included tentative proposals for ASL-4 evaluations and what their current best safety plan/argument for ASL-4 looks like (using just current science, no magic). Then, explain that plan wouldn’t be sufficient for reasonable amounts of safety (insofar as this is what they think).

Right now, they just have a bulleted list for ASL-4 countermeasures, but this is the main interesting thing at me. (I’m not really sold on substantial risk from systems which aren’t capable of carrying out that harm mostly autonomously, so I don’t think ASL-3 is actually important except as setup.)