Fetching the paper…
Reading the bibliography…
Do large language models (LLMs) have theory of mind? A plethora of papers and benchmarks have been introduced to evaluate if current models have been able to develop this key ability of social intelligence.
A formal basis for the heuristic determination of minimum cost paths
Peter E Hart, Nils J Nilsson, and Bertram Raphael · 1968
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff · 1978
Earlier work this paper cites.
Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception
Heinz Wimmer and Josef Perner · 1983
Earlier work this paper cites.
Theory-of-mind deficits and causal attributions
Peter Kinderman, Robin Dunbar, and Richard P Bentall · 1998
Earlier work this paper cites.
A new test of social sensitivity: Detection of faux pas in normal children and children with asperger syndrome
Simon Baron-Cohen, Michelle O’Riordan, Rosie Jones, Valerie Stone, and Kate Plaisted · 1999
Earlier work this paper cites.
The role of the orbitofrontal cortex in affective theory of mind deficits in criminal offenders with psychopathic tendencies
Simone G Shamay-Tsoory, Hagai Harari, Judith Aharon-Peretz, and Yechiel Levkovitz · 2010
Earlier work this paper cites.
Making minds: How theory of mind develops
Henry M Wellman · 2014
Earlier work this paper cites.
Machine theory of mind
Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, SM Ali Eslami, and Matthew Botvinick · 2018
Earlier work this paper cites.
Revisiting the evaluation of theory of mind through question answering
Matthew Le, Y-Lan Boureau, and Maximilian Nickel · 2019
Earlier work this paper cites.
Deepfake detection by analyzing convolutional traces
Luca Guarnera, Oliver Giudice, and Sebastiano Battiato · 2020
Earlier work this paper cites.
Cx-tom: Counterfactual explanations with theory-of-mind for enhancing human trust in image recognition models
Arjun R. Akula, Keze Wang, Changsong Liu, Sari Saba-Sadiya, Hongjing Lu, Sinisa Todorovic, Joyce Chai, and Song-Chun Zhu · 2021
Earlier work this paper cites.
Mindcraft: Theory of mind modeling for situated dialogue in collaborative tasks
Cristian-Paul Bara, CH-Wang Sky, and Joyce Chai · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Towards mutual theory of mind in human-ai interaction: How language reflects what students perceive about a virtual teaching assistant
Qiaosi Wang, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel · 2021
Earlier work this paper cites.
Fake it till you make it: face analysis in the wild using synthetic data alone
Erroll Wood, Tadas Baltrušaitis, Charlie Hewitt, Sebastian Dziadzio, Thomas J Cashman, and Jamie Shotton · 2021
Earlier work this paper cites.
Few-shot language coordination by modeling theory of mind
Hao Zhu, Graham Neubig, and Yonatan Bisk · 2021
Earlier work this paper cites.
Neural theory-of-mind? on the limits of social intelligence in large lms
Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi · 2022
Cited alongside, same era.
Symmetric machine theory of mind
Melanie Sclar, Graham Neubig, and Yonatan Bisk · 2022
Cited alongside, same era.
Symbolic knowledge distillation: from general language models to commonsense models
Peter West, Chandra Bhagavatula, Jack Hessel, Jena Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi · 2022
Cited alongside, same era.
Re3: Generating longer stories with recursive reprompting and revision
Kevin Yang, Yuandong Tian, Nanyun Peng, and Dan Klein · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Cited alongside, same era.
Multi 3 woz: A multilingual, multi-domain, multi-parallel dataset for training and evaluating culturally adapted task-oriented dialog systems
Tombench: Benchmarking theory of mind in large language models
Zhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu, Yunghwei Lai, Zexuan Xiong, et al · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Understanding social reasoning in language models with language models
Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah Goodman · 2024
Closest in time.
A notion of complexity for theory of mind via discrete world models
X. Angelo Huang, Emanuele La Malfa, Samuele Marro, Andrea Asperti, Anthony G. Cohn, and Michael J. Wooldridge · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Songbo Hu, Han Zhou, Mete Hergul, Milan Gritta, Guchun Zhang, Ignacio Iacobacci, Ivan Vulić, and Anna Korhonen · 2023
Cited alongside, same era.
Fantom: A benchmark for stress-testing machine theory of mind in interactions
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap · 2023
Cited alongside, same era.
Minding language models’ (lack of) theory of mind: A plug-and-play multi-character belief tracker
Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov · 2023
Cited alongside, same era.
How well do large language models perform on faux pas tests?
Natalie Shapira, Guy Zwirn, and Yoav Goldberg · 2023
Cited alongside, same era.
Large language models fail on trivial alterations to theory-of-mind tasks
Tomer Ullman · 2023
Cited alongside, same era.
Synthetic data, real errors: how (not) to publish and use synthetic data
Boris Van Breugel, Zhaozhi Qian, and Mihaela Van Der Schaar · 2023
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Closest in time.
MMToM-QA: Multimodal theory of mind question answering
Chuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua Tenenbaum, and Tianmin Shu · 2024
Closest in time.
Perceptions to beliefs: Exploring precursory inferences for theory of mind in large language models
Chani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim, Yeon Seonwoo, Yejin Choi, Alice Oh, and Hyunwoo Kim · 2024
Closest in time.
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, et al · 2024
Closest in time.
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang · 2024
Closest in time.
Source2synth: Synthetic data generation and curation grounded in real data sources
Alisia Lupidi, Carlos Gemmell, Nicola Cancedda, Jane Dwivedi-Yu, Jason Weston, Jakob Foerster, Roberta Raileanu, and Maria Lomeli · 2024
Closest in time.
https://openai.com/index/hello-gpt-4o
OpenAI, 2024 · 2024
Closest in time.
Muma-tom: Multi-modal multi-agent theory of mind
Haojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin, Layla Isik, Yen-Ling Kuo, and Tianmin Shu · 2024
Closest in time.
Testing theory of mind in large language models and humans
James WA Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, et al · 2024
Closest in time.
Tianlu Wang, Ilia Kulikov, Olga Golovneva, Ping Yu, Weizhe Yuan, Jane Dwivedi-Yu, Richard Yuanzhe Pang, Maryam Fazel-Zarandi, Jason Weston, and Xian Li · 2024
Closest in time.
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, and Yulan He · 2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng YU, Zhengying Liu, Yu Zhang, James Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu · 2024
Closest in time.