Strongly agree that causality is a useful and under-estimated perspective in AI discussions. I’ve found the literature on invariant causal prediction, helpful for understanding distributional generalization, especially some of the results in this paper. One of the papers you reference above puts the core issue well:
generalization guarantees must build on knowledge or assumptions on the “relatedness” of different training and testing domains
Based on the way I sometimes see the generalization of LLM or other models being discussed, its not clear to me that this idea has really been internalized by the more ML/engineering focused segments of the AI community.
As for lesswrong specifically, I think the prevalence of beliefs about LDT/FDT/newcomb-like problems impacts the way causality is sometimes discussed in an interesting way. I think both there would be great benefit in both wider AI circles and lesswrong in particular seeing increased discussion of causality.
Strongly agree that causality is a useful and under-estimated perspective in AI discussions. I’ve found the literature on invariant causal prediction, helpful for understanding distributional generalization, especially some of the results in this paper. One of the papers you reference above puts the core issue well:
Based on the way I sometimes see the generalization of LLM or other models being discussed, its not clear to me that this idea has really been internalized by the more ML/engineering focused segments of the AI community.
As for lesswrong specifically, I think the prevalence of beliefs about LDT/FDT/newcomb-like problems impacts the way causality is sometimes discussed in an interesting way. I think both there would be great benefit in both wider AI circles and lesswrong in particular seeing increased discussion of causality.