I’m struggling to understand how to think about reward. It sounds like if a hypothetical ML model does reward hacking or reward tampering, it would be because the training process selected for that behavior, not because the model is out to “get reward”; it wouldn’t be out to get anything at all. Is that correct?
I’m struggling to understand how to think about reward. It sounds like if a hypothetical ML model does reward hacking or reward tampering, it would be because the training process selected for that behavior, not because the model is out to “get reward”; it wouldn’t be out to get anything at all. Is that correct?