Megan Kinniment comments on Conditioning Generative Models

Megan Kinniment 26 Jun 2022 23:41 UTC
1 point
0
AF
For the newspaper and reddit post examples, I think false beliefs remain relevant since these are observations about beliefs. For example, the observation of BigCo announcing they have solved alignment is compatible with worlds where they actually have solved alignment, but also with worlds where BigCo have made some mistake and alignment hasn’t actually been solved, even though people in-universe believe that it has. These kinds of ‘mistaken alignment’ worlds seem like they would probably contaminate the conditioning to some degree at least. (Especially if there are ways that early deceptive AIs might be able to manipulate BigCo and others into making these kinds of mistakes).
- Adam Jermyn 27 Jun 2022 1:11 UTC
  LW: 1 AF: 1
  0
  AF Parent
  Fully agreed.