I’ve withdrawn the comment you were replying to on other grounds (see edit), but my response to this is somewhat similar to other commenters:
(In fairness, the two humans in the transcript also talk a decent amount in chained low-context platitudes, so some of this may be the humans’ fault. :P)
Yeah, that was the claim I was trying to make. I see you listing interpretations for how LaMDA could have come up with those responses without thinking very deeply. I don’t see you pointing out anything that a human clearly wouldn’t have done. I tend to assume that LaMDA does indeed make more egregiously nonhuman mistakes, like GPT also makes, but I don’t think we see them here.
I’m not particularly surprised if a human brings up meditation when asked about their inner contemplative life, even if the answer isn’t quite in the spirit of the question. Nor is an unexplained use of “kindred spirits” strikingly incoherent in that way.
Obviously, though, what we’re coming up against here is that it is pretty difficult/ambiguous to really decide what constitutes “human-level performance” here. Whether a given system “passes the Turing test” is incredibly dependent on the judge, and also, on which humans the system is competing with.
I’ve withdrawn the comment you were replying to on other grounds (see edit), but my response to this is somewhat similar to other commenters:
Yeah, that was the claim I was trying to make. I see you listing interpretations for how LaMDA could have come up with those responses without thinking very deeply. I don’t see you pointing out anything that a human clearly wouldn’t have done. I tend to assume that LaMDA does indeed make more egregiously nonhuman mistakes, like GPT also makes, but I don’t think we see them here.
I’m not particularly surprised if a human brings up meditation when asked about their inner contemplative life, even if the answer isn’t quite in the spirit of the question. Nor is an unexplained use of “kindred spirits” strikingly incoherent in that way.
Obviously, though, what we’re coming up against here is that it is pretty difficult/ambiguous to really decide what constitutes “human-level performance” here. Whether a given system “passes the Turing test” is incredibly dependent on the judge, and also, on which humans the system is competing with.