There is the issue of avoiding ignorant-yet-confident meta-preferences, which I’m working on writing up right now (partially thanks to you very comment here, thanks!)
I look forward to reading that. In the meantime can you address my parenthetical point in the grand-parent comment: “correctly extracting William MacAskill’s meta-preferences seems equivalent to learning metaphilosophy from William”? If it’s not clear, what I mean is that suppose Will wants to figure out his values by doing philosophy (which I think he actually does), does that mean that under you scheme the AI needs to learn how to do philosophy? If so, how do you plan to get around the problems with applying ML to metaphilosophy that I described in Some Thoughts on Metaphilosophy?
There is one way of doing metaphilosophy this way, which is “run (simulated) William MacAskill until he thinks he’s found a good metaphilosophy” or “find a description of metaphilosophy to which WA would say ‘yes’.”
But what the system I’ve sketched would most likely do is come up with something to which WA would say “yes, I can kinda see why that was built, but it doesn’t really fit together as I’d like and has a some of ad hoc and object level features”. That’s the “adequate” part of the process.
I look forward to reading that. In the meantime can you address my parenthetical point in the grand-parent comment: “correctly extracting William MacAskill’s meta-preferences seems equivalent to learning metaphilosophy from William”? If it’s not clear, what I mean is that suppose Will wants to figure out his values by doing philosophy (which I think he actually does), does that mean that under you scheme the AI needs to learn how to do philosophy? If so, how do you plan to get around the problems with applying ML to metaphilosophy that I described in Some Thoughts on Metaphilosophy?
There is one way of doing metaphilosophy this way, which is “run (simulated) William MacAskill until he thinks he’s found a good metaphilosophy” or “find a description of metaphilosophy to which WA would say ‘yes’.”
But what the system I’ve sketched would most likely do is come up with something to which WA would say “yes, I can kinda see why that was built, but it doesn’t really fit together as I’d like and has a some of ad hoc and object level features”. That’s the “adequate” part of the process.