I think this paper will be of interest. It’s a formal definition of universal intelligence/optimization power. Essentially you ask how well the agent does on average in an environment specified by a random program, where all rewards are specified by the environment program and observed by the agent. Unfortunately it’s uncomputable and requires a prior over environments.
I think this paper will be of interest. It’s a formal definition of universal intelligence/optimization power. Essentially you ask how well the agent does on average in an environment specified by a random program, where all rewards are specified by the environment program and observed by the agent. Unfortunately it’s uncomputable and requires a prior over environments.