This avoids spending lots of time getting confused about concepts that are confusing because they were the wrong thing to think about all along, such as “what is the shape of human values?” or “what does GPT4 want?”
These sound like exactly the sort of questions I’m most interested in answering. We live in a world of minds that have values and want things, and we are trying to prevent the creation of a mind that would be extremely dangerous to that world. These kind of questions feel to me like they tend to ground us to reality.
These sound like exactly the sort of questions I’m most interested in answering. We live in a world of minds that have values and want things, and we are trying to prevent the creation of a mind that would be extremely dangerous to that world. These kind of questions feel to me like they tend to ground us to reality.