AboutProjectsPublications

Theory of Mind in Large Language Models

GPT-4o Lacks Core Features of Theory of Mind

John Muchovej, Amanda Royka, Shane Lee, Julián Jara-Ettinger

Abstract

Do Large Language Models (LLMs) possess a Theory of Mind (ToM)? Research into this question has found that LLMs succeed on a range of benchmark tasks. However, these evaluations do not test for the actual representations posited by ToM: namely, a causal model of mental states and behavior. Here, we use a cognitively-grounded definition of ToM to develop and test a new evaluation framework. Specifically, our approach probes whether LLMs have a coherent, abstract, and consistent model of how mental states cause behavior – regardless of whether that model matches a human-like ToM. We test our evaluation against GPT-4o and find that even though it succeeds in approximating human judgments in a simple ToM paradigm, GPT-4o fails at a logically-equivalent task and exhibits low consistency between its action predictions and corresponding mental state inferences. As such, these findings suggest that GPT-4o’s social proficiency is not the result of a ToM.

Large Language Models Lack Core Features of Theory of Mind: Evidence from GPT-4o

John Muchovej, Shane Lee, Amanda Royka, Julián Jara-Ettinger

Abstract

Large Language Models (LLMs) have recently shown success across a range of social tasks, raising the question of whether they have a Theory of Mind (ToM). Research into this question has focused on evaluating LLMs against benchmarks, rather than testing for the representations posited by ToM. Using a cognitively-grounded definition of ToM, we develop a new evaluation framework that allows us to test whether LLMs have a mental causal model of other minds (ToM), human-like or not. We find that LLM social reasoning lacks key signatures expected from a causal model of other minds. These findings suggest that the social proficiency observed in LLMs is not the result of a ToM.