LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue

2026

Conference: Proceedings of SigDial

Katharine Kowalyshyn and Matthias Scheutz

Effective human teams excel at maintaining a consistent shared mental model (SMM) that reflects the shared understanding of individual team members about the task and what remains to be done. We present a novel, two-step framework that leverages large language models (LLMs) both as (1) generators of mental model traces from team dialogues to track the team’s SMM and (2) automated detectors of discrepancies between inferred and ground truth mental model traces. We define an SMM coherence evaluation framework for this use case and apply it to six dialogues in a previously published team corpus, ultimately producing a dataset of human and LLM SMM mental model traces, a reproducible evaluation framework for SMM coherence, and an empirical assessment of LLM-based discrepancy detection. Our results reveal that while LLMs exhibit apparent coherence on straightforward natural-language trace generation tasks, they systematically err in scenarios requiring spatial reasoning or disambiguation of transcription-level disfluencies.

@inproceedings{kowalyshynscheutz26sigdial,
  title={LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue},
  author={Katharine Kowalyshyn and Matthias Scheutz},
  year={2026},
  booktitle={Proceedings of SigDial},
  url={https://hrilab.tufts.edu/publications/kowalyshynscheutz26sigdial.pdf}
}