Joseph Bloom comments on Research Report: Sparse Autoencoders find only 9/180 board state features in OthelloGPT

Joseph Bloom 6 Mar 2024 0:35 UTC
4 points
0
@LawrenceC Nanda MATS stream played around with this as group project with code here: https://github.com/andyrdt/mats_sae_training/tree/othellogpt
- Robert_AIZI 6 Mar 2024 14:49 UTC
  2 points
  0
  Parent
  Cool! Do you know if they’ve written up results anywhere?
  - Joseph Bloom 6 Mar 2024 16:02 UTC
    3 points
    0
    Parent
    I think we got similar-ish results. @Andy Arditi was going to comment here to share them shortly.
    - Andy Arditi 6 Mar 2024 23:11 UTC
      3 points
      0
      Parent
      We haven’t written up our results yet.. but after seeing this post I don’t think we have to :P.
      
      We trained SAEs (with various expansion factors and L1 penalties) on the original Li et al model at layer 6, and found extremely similar results as presented in this analysis.
      
      It’s very nice to see independent efforts converge to the same findings!
      - Robert_AIZI 7 Mar 2024 15:00 UTC
        3 points
        0
        Parent
        Likewise, I’m glad to hear there was some confirmation from your team!
        
        An option for you if you don’t want to do a full writeup is to make a “diff” or comparison post, just listing where your methods and results were different (or the same). I think there’s demnad for that, people liked Comparing Anthropic’s Dictionary Learning to Ours
- LawrenceC 6 Mar 2024 0:44 UTC
  2 points
  0
  Parent
  Thanks!