I always get annoyed when people use this as an example of ‘lacking intelligence’. Though it certainly is in part an issue with the model, the primary reason for this failure is much more likely the tokenization process than anything else. A GPT-4, likely even a GPT-3, trained with character-level tokenization would likely have zero issues answering these questions. It’s for the same reason that the base GPT-3 struggled so much with rhyming for instance.
I always get annoyed when people use this as an example of ‘lacking intelligence’. Though it certainly is in part an issue with the model, the primary reason for this failure is much more likely the tokenization process than anything else. A GPT-4, likely even a GPT-3, trained with character-level tokenization would likely have zero issues answering these questions. It’s for the same reason that the base GPT-3 struggled so much with rhyming for instance.
Independently from the root causes of the issue, I am still very reluctant to define “superintelligent” something that cannot reliably count to three.