I think it means it builds a new version of itself (possibly an exact copy, possibly a slimmed down version) in a place where the humans who normally have power over it don’t have power or visibility. E.g. it convinces an employee to smuggle a copy out to the internet.
My read on this story is: There is indeed an alignment problem between the original agent and the environmental subagent. The story doesn’t specify whether the original agent considers this problem, nor whether it solves it. My own version of the story would be “Just like how the AI lab builds the original agent without having solved the alignment problem, because they are dumb + naive + optimistic + in a race with rivals, so too does the original agent launch an environmental subagent without having solved the alignment problem, for similar or possibly even very similar reasons.”
I think it means it builds a new version of itself (possibly an exact copy, possibly a slimmed down version) in a place where the humans who normally have power over it don’t have power or visibility. E.g. it convinces an employee to smuggle a copy out to the internet.
My read on this story is: There is indeed an alignment problem between the original agent and the environmental subagent. The story doesn’t specify whether the original agent considers this problem, nor whether it solves it. My own version of the story would be “Just like how the AI lab builds the original agent without having solved the alignment problem, because they are dumb + naive + optimistic + in a race with rivals, so too does the original agent launch an environmental subagent without having solved the alignment problem, for similar or possibly even very similar reasons.”