Signer comments on New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”

Signer 7 Dec 2023 8:12 UTC
LW: 3 AF: 3
0
AF
Wait, where? I think the objection to “Doing that is quite hard” is not an objection to “it’s not obviously true that such algorithms are actually “achievable” for SGD”—it’s an objection to the conclusion that model would try hard enough to justify arguments about deception from weak statement about loss decreasing during training.
- TurnTrout 26 Dec 2023 20:30 UTC
  LW: 2 AF: 2
  0
  AF Parent
  an objection to the conclusion that model would try hard enough to justify arguments about deception from weak statement about loss decreasing during training.
  This is… roughly one point I was making, yes.