gwern comments on EfficientZero: human ALE sample-efficiency w/MuZero+self-supervised

gwern 4 Nov 2021 16:14 UTC
LW: 12 AF: 9
AF
Speaking of sample-efficient MuZero, DM just posted “Procedural Generalization by Planning with Self-Supervised World Models”, Anand et al 2021 on a different set of environments than ALE, which also uses self-supervision to boost MuZero-Reanalyze sample-efficiency dramatically (in addition to another one of my pet interests, demonstrating implicit meta-learning through the blessings of scale): https://arxiv.org/pdf/2111.01587.pdf#page=5

Eyeballing the graph at 200k env frames, the self-supervised variant is >200x more sample-efficient than PPO (which has no sample-efficient variant the way DQN does because it’s on-policy and can’t reuse data), and 5x the baseline MuZero (and a ‘model-free’ MuZero variant/ablation I’m unfamiliar with). So reasonably parallel.
What links here?
- gwern's comment on Interpretability by abergal (7 Nov 2021 15:45 UTC; 3 points)
- John Schulman 7 Nov 2021 10:56 UTC
  LW: 4 AF: 3
  AF Parent
  Performance is mostly limited here by the fact that there are 500 levels for each game (i.e., level overfitting is the problem) so it’s not that meaningful to look at sample efficiency wrt environment interactions. The results would look a lot different on the full distribution of levels. I agree with your statement directionally though.
  - Ankesh Anand 7 Nov 2021 20:06 UTC
    LW: 9 AF: 6
    AF Parent
    We do actually train/evaluate on the full distribution (See Figure 5 rightmost). MuZero+SSL versions (especially reconstruction) continue to be a lot more sample-efficient even in the full-distribution, and MuZero itself seems to be quite a bit more sample efficient than PPO/PPG.
    - John Schulman 12 Nov 2021 10:06 UTC
      LW: 4 AF: 4
      AF Parent
      I’m still not sure how to reconcile your results with the fact that the participants in the procgen contest ended up winning with modifications of our PPO/PPG baselines, rather than Q-learning and other value-based algorithms, whereas your paper suggests that Q-learning performs much better. The contest used 8M timesteps + 200 levels. I assume that your “QL” baseline is pretty similar to widespread DQN implementations.
      https://arxiv.org/pdf/2103.15332.pdf
      https://www.aicrowd.com/challenges/neurips-2020-procgen-competition/leaderboards?challenge_leaderboard_extra_id=470&challenge_round_id=662
      Are there implementation level changes that dramatically improve performance of your QL implementation?
      (Currently on vacation and I read your paper briefly while traveling, but I may very well have missed something.)
      - Ankesh Anand 14 Nov 2021 3:45 UTC
        LW: 6 AF: 5
        AF Parent
        The Q-Learning baseline is a model-free control of MuZero. So it shares implementation details of MuZero (network architecture, replay ratio, training details etc.) while removing the model-based components of MuZero (details in sec A.2) . Some key differences you’d find vs a typical Q-learning implementation:
        Larger network architectures: 10 block ResNet compared to a few conv layers in typical implementations.
        Higher sample reuse: When using a reanalyse ratio of 0.95, both MuZero and Q-Learning use each replay buffer sample an average of 20 times. The target network is updated every 100 training steps.
        Batch size of 1024 and some smaller details like using categorical reward and value predictions similar to MuZero.
        We also have a small model-based component which predicts reward at next time step which lets us decompose the Q(s,a) into reward and value predictions just like MuZero.
        I would guess larger networks + higher sample reuse have the biggest effect size compared to standard Q-learning implementations.
        The ProcGen competition also might have used the easy difficulty mode compared to the hard difficulty mode used in our paper.
        John Schulman 14 Nov 2021 18:02 UTC
        LW: 3 AF: 3
        AF Parent
        Thanks, this is very insightful. BTW, I think your paper is excellent!
        Ankesh Anand 14 Nov 2021 21:26 UTC
        1 point
        AF Parent
        Thanks, glad you liked it, I really like the recent RL directions from OpenAI too! It would be interesting to see the use of model-based RL for the “RL as fine-tuning paradigm”: making large pre-trained models more aligned/goal-directed efficiently by simply searching over a reward function learned from humans.
        John Schulman 19 Nov 2021 8:59 UTC
        LW: 3 AF: 2
        AF Parent
        Would you say Learning to Summarize is an example of this? https://arxiv.org/abs/2009.01325
        It’s model based RL because you’re optimizing against the model of the human (ie the reward model). And there are some results at the end on test-time search.
        Or do you have something else in mind?
    - John Schulman 12 Nov 2021 9:58 UTC
      LW: 2 AF: 2
      AF Parent
      There’s no PPO/PPG curve there—I’d be curious to see that comparison. (though I agree that QL/MuZero will probably be more sample efficient.)
      - Ankesh Anand 14 Nov 2021 4:10 UTC
        2 points
        Parent
        I was eyeballing Figure 2 in the PPG paper and comparing it to our results on the full distribution (Table A.3).
        PPO: ~0.25
        PPG: ~0.52
        MuZero: 0.68
        MuZero+Reconstruction: 0.93

gwern comments on EfficientZero: human ALE sample-efficiency w/​MuZero+self-supervised

gwern comments on EfficientZero: human ALE sample-efficiency w/MuZero+self-supervised