So the definition of myopia given in Defining Myopia was quite similar to my expansion in the But Wait There’s More section; you can roughly match them up by saying r(x)=∑ifiri(x) and yi(x)=(1−fi)ri(x) , where fi is a real number corresponding to the amount that the agent cares about rewards obtained in episode i and ri is the reward obtained in episode i. Putting both of these into the sum gives R(x)=∑iri(x), the undiscounted, non-myopic reward that the agent eventually obtains.
In terms of the R=R0+R1 definition that I give in the uncertainty framing, this is R0=R(x,y0)=∑ifiri(x)+∑i(1−fi)ri(x0), and R1=R(x,y)−R(x,y0)=∑i(1−fi)(ri(x)−ri(x0)).
So if you let r be a vector of the reward obtained on each step and f be a vector of how much the agent cares about each step then x→x+ϵ∑ifi∂ri∂x , and thus the change to the overall reward is R→R+ϵ∑i∂ri∂x∑jfj∂rj∂x , which can be negative if the two sums have different signs.
I was hoping that a point would reveal itself to me about now but I’ll have to get back to you on that one.
So the definition of myopia given in Defining Myopia was quite similar to my expansion in the But Wait There’s More section; you can roughly match them up by saying r(x)=∑ifiri(x) and yi(x)=(1−fi)ri(x) , where fi is a real number corresponding to the amount that the agent cares about rewards obtained in episode i and ri is the reward obtained in episode i. Putting both of these into the sum gives R(x)=∑iri(x), the undiscounted, non-myopic reward that the agent eventually obtains.
In terms of the R=R0+R1 definition that I give in the uncertainty framing, this is R0=R(x,y0)=∑ifiri(x)+∑i(1−fi)ri(x0), and R1=R(x,y)−R(x,y0)=∑i(1−fi)(ri(x)−ri(x0)).
So if you let r be a vector of the reward obtained on each step and f be a vector of how much the agent cares about each step then x→x+ϵ∑ifi∂ri∂x , and thus the change to the overall reward is R→R+ϵ∑i∂ri∂x∑jfj∂rj∂x , which can be negative if the two sums have different signs.
I was hoping that a point would reveal itself to me about now but I’ll have to get back to you on that one.