JesseClifton comments on The Commitment Races problem

JesseClifton 1 Mar 2021 0:09 UTC
LW: 6 AF: 3
AF
It seems like we can kind of separate the problem of equilibrium selection from the problem of “thinking more”, if “thinking more” just means refining one’s world models and credences over them. One can make conditional commitments of the form: “When I encounter future bargaining partners, we will (based on our models at that time) agree on a world-model according to some protocol and apply some solution concept (e.g. Nash or Kalai-Smorodinsky) to it in order to arrive at an agreement.”

The set of solution concepts you commit to regarding as acceptable still poses an equilibrium selection problem. But, on the face of it at least, the “thinking more” part is handled by conditional commitments to act on the basis of future beliefs.

I guess there’s the problem of what protocols for specifying future world-models you commit to regarding as acceptable. Maybe there are additional protocols that haven’t occurred to you, but which other agents may have committed to and which you would regard as acceptable when presented to you. Hopefully it is possible to specify sufficiently flexible methods for determining whether protocols proposed by your future counterparts are acceptable that this is not a problem.
- Daniel Kokotajlo 1 Mar 2021 17:47 UTC
  LW: 3 AF: 2
  AF Parent
  If I read you correctly, you are suggesting that some portion of the problem can be solved, basically—that it’s in some sense obviously a good idea to make a certain sort of commitment, e.g. “When I encounter future bargaining partners, we will (based on our models at that time) agree on a world-model according to some protocol and apply some solution concept (e.g. Nash or Kalai-Smorodinsky) to it in order to arrive at an agreement.” So the commitment races problem may still exist, but it’s about what other commitments to make besides this one, and when. Is this a fair summary?
  I guess my response would be “On the object level, this seems like maybe a reasonable commitment to me, though I’d have lots of questions about the details. We want it to be vague/general/flexible enough that we can get along nicely with various future agents with somewhat different protocols, and what about agents that are otherwise reasonable and cooperative but for some reason don’t want to agree on a world-model with us? On the meta level though, I’m still feeling burned from the various things that seemed like good commitments to me and turned out to be dangerous, so I’d like to have some sort of stronger reason to think this is safe.”
  - JesseClifton 1 Mar 2021 19:46 UTC
    LW: 3 AF: 2
    AF Parent
    Yeah I agree the details aren’t clear. Hopefully your conditional commitment can be made flexible enough that it leaves you open to being convinced by agents who have good reasons for refusing to do this world-model agreement thing. It’s certainly not clear to me how one could do this. If you had some trusted “deliberation module”, which engages in open-ended generation and scrutiny of arguments, then maybe you could make a commitment of the form “use this protocol, unless my counterpart provides reasons which cause my deliberation module to be convinced otherwise”. Idk.
    
    Your meta-level concern seems warranted. One would at least want to try to formalize the kinds of commitments we’re discussing and ask if they provide any guarantees, modulo equilibrium selection.
    - Daniel Kokotajlo 1 Mar 2021 22:14 UTC
      LW: 2 AF: 1
      AF Parent
      I think we are on the same page then. I like the idea of a deliberation module; it seems similar to the “moral reasoning module” I suggested a while back. The key is to make it not itself a coward or bully, reasoning about schelling points and universal principles and the like instead of about what-will-lead-to-the-best-expected-outcomes-given-my-current-credences.