beren comments on The Computational Anatomy of Human Values

beren 8 Apr 2023 11:09 UTC
3 points
0
So, I agree and I think we are getting at the same thing (though not completely sure what you are pointing at). The way to have a model-y critic and actor is to have the actor and critic perform model-free RL over the latent space of your unsupervised world model. This is the key point of my post and why humans can have ‘values’ and desires for highly abstract linguistic concepts such as ‘justice’ as opposed to pure sensory states or primary rewards.