I think Redwood’s classifier project was a reasonable project to work towards, and I think this post was great because it both displayed a bunch of important virtues and avoided doubling down on trying to always frame one’s research in a positive light.
I was really very glad to see this update come out at the time, and it made me hopeful that we can have a great discourse on LessWrong and AI Alignment where when people sometimes overstate things, they can say “oops”, learn and move on. My sense is Redwood made a pretty deep update from the first post they published (and this update), and hasn’t made any similar errors since then.
I think Redwood’s classifier project was a reasonable project to work towards, and I think this post was great because it both displayed a bunch of important virtues and avoided doubling down on trying to always frame one’s research in a positive light.
I was really very glad to see this update come out at the time, and it made me hopeful that we can have a great discourse on LessWrong and AI Alignment where when people sometimes overstate things, they can say “oops”, learn and move on. My sense is Redwood made a pretty deep update from the first post they published (and this update), and hasn’t made any similar errors since then.