Robert_AIZI comments on Research Report: Sparse Autoencoders find only 9/180 board state features in OthelloGPT

Robert_AIZI 5 Mar 2024 18:44 UTC
5 points
0
Thanks for doing this—could you share your code?
Just uploaded the code here: https://github.com/RobertHuben/othellogpt_sparse_autoencoders/. Apologies in advance, the code is kinda a mess since I’ve been writing for myself. I’ll take the hour or so to add info to the readme about the files and how to replicate my experiments.
I’d be interested in running the code on the model used by Li et al, which he’s hosted on Google drive:
https://drive.google.com/drive/folders/1bpnwJnccpr9W-N_hzXSm59hT7Lij4HxZ
Thanks for the link! I think replicating the experiment with Li et al’s model is definite next step! Perhaps we can have a friendly competition to see who writes it up first :)

I have mixed feelings about whether the results will be different with the high-accuracy model from Li et al:
1. On priors, if the features are more “unambiguous”, they should be easier for the sparse autoencoder to find.
2. But my hacky model was at least trained enough that those features do emerge from linear probes. If sparse autoencoders can’t match linear probes, thats also worth knowing.
3. If there is a difference, and sparse autoencoders work on a language model that’s sufficiently trained, would LLMs meet that criteria?
Also, in addition to the future work you list, I’d be interested in running the SAEs with much larger Rs and with alternative hyperparameter selection criteria.
Agree that its worth experimenting with R, but the only other hyperparameter is the sparsity coefficient alpha, and I found that alpha had to be in a narrow range or the training would collapse to “all variance is unexplained” or “no active features”. (Maybe you mean Adam hyperparameters, which I suppose might also be worth experimenting with.) Here’s the result of my hyperparameter sweep for alpha:
- leogao 5 Mar 2024 23:51 UTC
  4 points
  2
  Parent
  Fwiw, I find it’s much more useful to have (log) active features on the x axis, and (log) unexplained variance on the y axis. (if you want you can then also plot the L1 coefficient above the points, but that seems less important)
  - Robert_AIZI 6 Mar 2024 15:05 UTC
    2 points
    0
    Parent
    Good thinking, here’s that graph! I also annotated it to show where the alpha value I ended up using for the experiment. Its improved over the pareto frontier shown on the graph, and I believe thats because the data in this sweep was from training for 1 epoch, and the real run I used for the SAE was 4 epochs.
    - leogao 7 Mar 2024 0:20 UTC
      5 points
      0
      Parent
      In my experiments log L0 vs log unexplained variance should be a nice straight line. I think your autoencoders might be substantially undertrained (especially given that training longer moves off the frontier a lot). Scaling up the data by 10x or 100x wouldn’t be crazy.
      (Also, I think L0 is more meaningful than L0 / d_hidden for comparing across different d_hidden (I assume that’s what “percent active features” is))
- LawrenceC 5 Mar 2024 19:31 UTC
  2 points
  0
  Parent
  Thanks for uploading your interp and training code!
  Could you upload your model and/or datasets somewhere as well, for reproducibility? (i.e. your datasets folder containing the datasets:)
```
def recognized_dataset():
    mode_lookups={
        "gpt_train":        ["datasets/othello_gpt_training_corpus.txt",        OthelloDataset,         {}],
        "gpt_train_small":  ["datasets/small_othello_gpt_training_corpus.txt",  OthelloDataset,         {}],
        "gpt_test":         ["datasets/othello_gpt_test_corpus.txt",            OthelloDataset,         {}],
        "sae_train":        ["datasets/sae_training_corpus.txt",                OthelloDataset,         {}],
        "probe_train":      ["datasets/probe_train_corpus.txt",                 LabelledOthelloDataset, {}],
        "probe_train_bw":   ["datasets/probe_train_corpus.txt",                 LabelledOthelloDataset, {"use_ally_enemy":False}],
        "probe_train_small":["datasets/small_probe_training_corpus.txt",        LabelledOthelloDataset, {}],
        "probe_test":       ["datasets/probe_test_corpus.txt",                  LabelledOthelloDataset, {}],
    }
    return mode_lookups
```
  Agree that its worth experimenting with R, but the only other hyperparameter is the sparsity coefficient alpha, and I found that alpha had to be in a narrow range or the training would collapse to “all variance is unexplained” or “no active features”.
  Yeah, the main hyperparameters are the expansion factor and “what optimization algorithm do you use/what hyperparameters do you use for the optimization algorithm”.
  - Robert_AIZI 5 Mar 2024 19:58 UTC
    2 points
    0
    Parent
    Here are the datasets, OthelloGPT model (“trained_model_full.pkl”), autoencoders (saes/), probes, and a lot of the cached results (it takes a while to compute AUROC for all position/feature pairs, so I found it easier to save those): https://drive.google.com/drive/folders/1CSzsq_mlNqRwwXNN50UOcK8sfbpU74MV
    
    You should download all of these into the same level directory as the main repo.
    - Joseph Bloom 6 Mar 2024 0:35 UTC
      4 points
      0
      Parent
      @LawrenceC Nanda MATS stream played around with this as group project with code here: https://github.com/andyrdt/mats_sae_training/tree/othellogpt
      - Robert_AIZI 6 Mar 2024 14:49 UTC
        2 points
        0
        Parent
        Cool! Do you know if they’ve written up results anywhere?
        Joseph Bloom 6 Mar 2024 16:02 UTC
        3 points
        0
        Parent
        I think we got similar-ish results. @Andy Arditi was going to comment here to share them shortly.
        Andy Arditi 6 Mar 2024 23:11 UTC
        3 points
        0
        Parent
        We haven’t written up our results yet.. but after seeing this post I don’t think we have to :P.
        
        We trained SAEs (with various expansion factors and L1 penalties) on the original Li et al model at layer 6, and found extremely similar results as presented in this analysis.
        
        It’s very nice to see independent efforts converge to the same findings!
        Robert_AIZI 7 Mar 2024 15:00 UTC
        3 points
        0
        Parent
        Likewise, I’m glad to hear there was some confirmation from your team!
        
        An option for you if you don’t want to do a full writeup is to make a “diff” or comparison post, just listing where your methods and results were different (or the same). I think there’s demnad for that, people liked Comparing Anthropic’s Dictionary Learning to Ours
      - LawrenceC 6 Mar 2024 0:44 UTC
        2 points
        0
        Parent
        Thanks!

Robert_AIZI comments on Research Report: Sparse Autoencoders find only 9/​180 board state features in OthelloGPT

Robert_AIZI comments on Research Report: Sparse Autoencoders find only 9/180 board state features in OthelloGPT