Fetching the paper…
Reading the bibliography…
Existing LLM reasoning methods have shown impressive capabilities across various tasks, such as solving math and coding problems.
The Intentional Stance
Daniel Clement Dennett · 1981
Earlier work this paper cites.
Goal attribution without agency cues: the perception of ‘pure reason’in infancy
Gergely Csibra, Gyórgy Gergely, Szilvia Bíró, Orsolya Koos, and Margaret Brockbank · 1999
Earlier work this paper cites.
A sequential particle filter method for static models
Nicolas Chopin · 2002
Earlier work this paper cites.
Comparison of resampling schemes for particle filtering
Randal Douc and Olivier Cappé · 2005
Earlier work this paper cites.
Sequential monte carlo samplers
Pierre Del Moral, Arnaud Doucet, and Ajay Jasra · 2006
Earlier work this paper cites.
Action understanding as inverse planning
Chris L Baker, Rebecca Saxe, and Joshua B Tenenbaum · 2009
Earlier work this paper cites.
Monte carlo sampling methods for approximating interactive pomdps
Prashant Doshi and Piotr J Gmytrasiewicz · 2009
Earlier work this paper cites.
Data-driven sequential monte carlo in probabilistic programming
Yura N Perov, Tuan Anh Le, and Frank Wood · 2015
Earlier work this paper cites.
Psychological reasoning in infancy
Renée Baillargeon, Rose M Scott, and Lin Bian · 2016
Earlier work this paper cites.
Rational quantitative attribution of beliefs, desires and percepts in human mentalizing
Chris L. Baker, Julian Jara-Ettinger, Rebecca Saxe, and Joshua B. Tenenbaum · 2017
Earlier work this paper cites.
The naive utility calculus as a unified, quantitative framework for action understanding
Julian Jara-Ettinger, Laura Schulz, and Josh Tenenbaum · 2019
Earlier work this paper cites.
Revisiting the evaluation of theory of mind through question answering
Matthew Le, Y-Lan Boureau, and Maximilian Nickel · 2019
Earlier work this paper cites.
Online bayesian goal inference for boundedly rational planning agents
Tan Zhi-Xuan, Jordyn Mann, Tom Silver, Josh Tenenbaum, and Vikash Mansinghka · 2020
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Neural theory-of-mind? on the limits of social intelligence in large LMs
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Solving the baby intuitions benchmark with a hierarchically bayesian theory of mind
Tan Zhi-Xuan, Nishad Gothoskar, Falk Pollok, Dan Gutfreund, Joshua B Tenenbaum, and Vikash K Mansinghka · 2022
Earlier work this paper cites.
A (dis-) information theory of revealed and unrevealed preferences: emerging deception and skepticism via theory of mind
Nitay Alon, Lion Schulz, Jeffrey S Rosenschein, and Peter Dayan · 2023
Earlier work this paper cites.
Understanding social reasoning in language models with language models
Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah Goodman · 2023
Cited alongside, same era.
FANToM: A benchmark for stress-testing machine theory of mind in interactions
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap · 2023
Cited alongside, same era.
Towards a holistic landscape of situated theory of mind in large language models
Ziqiao Ma, Jacob Sansom, Run Peng, and Joyce Chai · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Cited alongside, same era.
Minding language models’ (lack of) theory of mind: A plug-and-play multi-character belief tracker
Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov · 2023
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini-Team · 2024
Later among the works it cites.
A notion of complexity for theory of mind via discrete world models
X Angelo Huang, Emanuele La Malfa, Samuele Marro, Andrea Asperti, Anthony Cohn, and Michael Wooldridge · 2024
Later among the works it cites.
MMToM-QA: Multimodal theory of mind question answering
Chuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua Tenenbaum, and Tianmin Shu · 2024
Later among the works it cites.
Perceptions to beliefs: Exploring precursory inferences for theory of mind in large language models
Chani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim, Yeon Seonwoo, Yejin Choi, Alice Oh, and Hyunwoo Kim · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How well do large language models perform on faux pas tests?
Natalie Shapira, Guy Zwirn, and Yoav Goldberg · 2023
Cited alongside, same era.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao · 2023
Cited alongside, same era.
Large language models fail on trivial alterations to theory-of-mind tasks
Tomer Ullman · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Cited alongside, same era.
Can a machine know that we know what it knows?
Oliver Whang · 2023
Cited alongside, same era.
Hi-ToM: A benchmark for evaluating higher-order theory of mind reasoning in large language models
Yufan Wu, Yinghui He, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik R Narasimhan · 2023
Cited alongside, same era.
Aliakbar Nafar, Kristen Brent Venable, and Parisa Kordjamshidi · 2024
Later among the works it cites.
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Qwen-Team · 2024
Later among the works it cites.
Clever hans or neural theory of mind? stress testing social reasoning in large language models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz · 2024
Later among the works it cites.
Think twice: Perspective-taking improves large language models’ theory-of-mind capabilities
Alex Wilf, Sihyun Lee, Paul Pu Liang, and Louis-Philippe Morency · 2024
Later among the works it cites.
OpenToM: A comprehensive benchmark for evaluating theory-of-mind reasoning capabilities of large language models
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, and Yulan He · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Probabilistic inference in language models via twisted sequential monte carlo
Stephen Zhao, Rob Brekelmans, Alireza Makhzani, and Roger Baker Grosse · 2024
Later among the works it cites.
Hypothetical minds: Scaffolding theory of mind for multi-agent tasks with large language models
Logan Cross, Violet Xiang, Agam Bhatia, Daniel LK Yamins, and Nick Haber · 2025
Closest in time.
Syntactic and semantic control of large language models via sequential monte carlo
João Loula, Benjamin LeBrun, Li Du, Ben Lipkin, Clemente Pasti, Gabriel Grand, Tianyu Liu, Yahya Emara, Marjorie Freedman, Jason Eisner, Ryan Cotterell, Vikash Mansinghka, Alexander K. Lew, Tim Vieira, and Timothy J. O’Donnell · 2025
Closest in time.
o3-mini system card, 2025
OpenAI · 2025
Closest in time.
Explore theory of mind: Program-guided adversarial data generation for theory of mind reasoning
Melanie Sclar, Jane Dwivedi-Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov, Yonatan Bisk, Yejin Choi, and Asli Celikyilmaz · 2025
Closest in time.
Thoughts are all over the place: On the underthinking of o1-like llms
Yue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Linfeng Song, Dian Yu, Juntao Li, Zhuosheng Zhang, et al · 2025
Closest in time.
Understanding epistemic language with a language-augmented bayesian theory of mind
Lance Ying, Tan Zhi-Xuan, Lionel Wong, Vikash Mansinghka, and Joshua B Tenenbaum · 2025
Closest in time.
Pragmatic instruction following and goal assistance via cooperative language-guided inverse planning
Tan Zhi-Xuan, Lance Ying, Vikash Mansinghka, and Joshua B Tenenbaum · 2094
Closest in time.