If the LLM says “yes”, then tell it “That makes sense! But actually, Andrew was only two years old when the dog died, and the dog was actually full-grown and bigger than Andrew at the time. Do you still think Andrew was able to lift up the dog?”, and it will probably say “no”. Then say “That makes sense as well. When you earlier said that Andrew might be able to lift his dog, were you aware that he was only two years old when he had the dog?” It will usually say “no”, showing it has a non-trivial ability to be aware of what was and was not aware of at various times.
This doesn’t demonstrate anything about awareness of awareness. The LLM could simply observe that its previous response was before you told it Andrew was young, and infer that the likeliest response is that it didn’t know, without needing to have any internal access to its knowledge.
This doesn’t demonstrate anything about awareness of awareness. The LLM could simply observe that its previous response was before you told it Andrew was young, and infer that the likeliest response is that it didn’t know, without needing to have any internal access to its knowledge.