Did you know about “by default, GPTs think in plain sight”? It doesn’t explicitly talk about agentized GPTs but was discussing the impact this has on GPTs for AGI and how it affects the risks, and what we should do about it (eg. maybe rlhf is dangerous)
Thank you. I think it is relevant. I just found it yesterday following up on this. The comment there by Gwern is a really interesting example of how we could accidentally introduce pressure for them to use steganography so their thoughts aren’t in English.
What I’m excited about is that agentizing them, while dangerous, could mean they not only think in plain sight, but they’re actually what gets used. That would cross from only being able to say how to get alignment, to making it so.ething the world would actually do.
Did you know about “by default, GPTs think in plain sight”?
It doesn’t explicitly talk about agentized GPTs but was discussing the impact this has on GPTs for AGI and how it affects the risks, and what we should do about it (eg. maybe rlhf is dangerous)
Thank you. I think it is relevant. I just found it yesterday following up on this. The comment there by Gwern is a really interesting example of how we could accidentally introduce pressure for them to use steganography so their thoughts aren’t in English.
What I’m excited about is that agentizing them, while dangerous, could mean they not only think in plain sight, but they’re actually what gets used. That would cross from only being able to say how to get alignment, to making it so.ething the world would actually do.