Or something simpler would be that the agent’s money counter is in the environment but unmodifiable except by getting tokens, and the agent’s goal is to maximize this quantity. Feels kind of fake maybe because money gives the agent no power or intelligence, but it’s a valid object-in-the-world to have a preference over the state of.
Yet another option is to have the agent maximize energy tokens (which actions consume)
Or something simpler would be that the agent’s money counter is in the environment but unmodifiable except by getting tokens, and the agent’s goal is to maximize this quantity. Feels kind of fake maybe because money gives the agent no power or intelligence, but it’s a valid object-in-the-world to have a preference over the state of.
Yet another option is to have the agent maximize energy tokens (which actions consume)