I’m in the process of trying to build an org focused on “automated/augmented alignment research.” As part of that, I’ve been thinking about which alignment research agendas could be investigated in order to make automated alignment safer and trustworthy. And so, I’ve been thinking of doing internal research on AI control/security and using that research internally to build parts of the system I intend to build. I figured this would be a useful test case for applying the AI control agenda and iterating on issues we face in implementation, and then sharing those insights with the wider community.
Would love to talk to anyone who has thoughts on this or who would introduce me to someone who would fund this kind of work.
I’m in the process of trying to build an org focused on “automated/augmented alignment research.” As part of that, I’ve been thinking about which alignment research agendas could be investigated in order to make automated alignment safer and trustworthy. And so, I’ve been thinking of doing internal research on AI control/security and using that research internally to build parts of the system I intend to build. I figured this would be a useful test case for applying the AI control agenda and iterating on issues we face in implementation, and then sharing those insights with the wider community.
Would love to talk to anyone who has thoughts on this or who would introduce me to someone who would fund this kind of work.