I think that is exactly right.
I also wouldn’t be too surprised if in some domains RL leads to useful agents if all the individual actions are known to and doable by the model and RL teaches it how to sensibly string these actions together. This doesn’t seem too different from mathematical derivations.
I kinda agree with this as well. Except that it seems completely unclear to me whether recreating the missing human capabilities/brain systems takes two years or two decades or even longer.
It doesn’t seem to me to be a single missing thing and for each separate step holds: That it hasn’t been done yet is evidence that it’s not that easy.