I notice that I don’t expect FOOM like RSI, because I don’t expect we’ll get an mesa optimizer with coherent goals. It’s not hard to give the outer optimiser (e.g. gradient decent) a coherent goal. For the outer optimiser to have a coherent goal is the default. But I don’t expect that to translate to the inner optimiser. The inner optimiser will just have a bunch of heuristics and proxi-goals, and not be very coherent, just like humans.
The outer optimiser can’t FOOM, since it don’t do planing, and don’t have strategic self awareness. It’s can only do some combination of hill climbing and random trial and error. If something is FOOMing it will be the inner optimiser, but I expect that one to be a mess.
I notice that this argument don’t quite hold. More coherence is useful for RSI, but complete coherence is not necessary.
I also notice that I expect AIs to make fragile plans, but on reflection, I expect them to gett better and better with this. By fragile I mean that the longer the plan is, the more likely it is to break. This is true for human too though. But we are self aware enough about this fact to mostly compensate, i.e. make plans that don’t have too many complicated steps, even if the plan spans a long time.
Recording though in progress...
I notice that I don’t expect FOOM like RSI, because I don’t expect we’ll get an mesa optimizer with coherent goals. It’s not hard to give the outer optimiser (e.g. gradient decent) a coherent goal. For the outer optimiser to have a coherent goal is the default. But I don’t expect that to translate to the inner optimiser. The inner optimiser will just have a bunch of heuristics and proxi-goals, and not be very coherent, just like humans.
The outer optimiser can’t FOOM, since it don’t do planing, and don’t have strategic self awareness. It’s can only do some combination of hill climbing and random trial and error. If something is FOOMing it will be the inner optimiser, but I expect that one to be a mess.
I notice that this argument don’t quite hold. More coherence is useful for RSI, but complete coherence is not necessary.
I also notice that I expect AIs to make fragile plans, but on reflection, I expect them to gett better and better with this. By fragile I mean that the longer the plan is, the more likely it is to break. This is true for human too though. But we are self aware enough about this fact to mostly compensate, i.e. make plans that don’t have too many complicated steps, even if the plan spans a long time.