johnswentworth comments on Ngo and Yudkowsky on alignment difficulty

johnswentworth 19 Nov 2021 16:54 UTC
LW: 21 AF: 13
AF
I do think alignment has a relatively-simple core. Not as simple as intelligence/competence, since there’s a decent number of human-value-specific bits which need to be hardcoded (as they are in humans), but not enough to drive the bulk of the asymmetry.
(BTW, I do think you’ve correctly identified an important point which I think a lot of people miss: humans internally “learn” values from a relatively-small chunk of hardcoded information. It should be possible in-principle to specify values with a relatively small set of hardcoded info, similar to the way humans do it; I’d guess fewer than at most 1000 things on the order of complexity of a very fuzzy face detector are required, and probably fewer than 100.)
The reason it’s less learnable than competence is not that alignment is much more complex, but that it’s harder to generate a robust reward signal for alignment. Basically any sufficiently-complex long-term reward signal should incentivize competence. But the vast majority of reward signals do not incentivize alignment. In particular, even if we have a reward signal which is “close” to incentivizing alignment in some sense, the actual-process-which-generates-the-reward-signal is likely to be at least as simple/natural as actual alignment.
(I’ll note that the departure from talking about Hidden Complexity here is mainly because competence in particular is a special case where “complexity” plays almost no role, since it’s incentivized by almost any reward. Hidden Complexity is still usually the right tool for talking about why any particular reward-signal will not incentivize alignment.)
I suspect that Eliezer’s answer to this would be different, and I don’t have a good guess what it would be.
- cousin_it 22 Nov 2021 17:32 UTC
  LW: 30 AF: 15
  AF Parent
  Thinking about it more, it seems that messy reward signals will lead to some approximation of alignment that works while the agent has low power compared to its “teachers”, but at high power it will do something strange and maybe harm the “teachers” values. That holds true for humans gaining a lot of power and going against evolutionary values (“superstimuli”), and for individual humans gaining a lot of power and going against societal values (“power corrupts”), so it’s probably true for AI as well. The worrying thing is that high power by itself seems sufficient for the change, for example if an AI gets good at real-world planning, that constitutes power and therefore danger. And there don’t seem to be any natural counterexamples. So yeah, I’m updating toward your view on this.