Fetching the paper…
Reading the bibliography…
Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity.
Child’s Conception of Space
Jean Piaget. 1956 · 1956
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff. 1978 · 1978
Earlier work this paper cites.
Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception
Heinz Wimmer and Josef Perner. 1983 · 1983
Earlier work this paper cites.
Does the autistic child have a “theory of mind”?
Simon Baron-Cohen, Alan M Leslie, and Uta Frith. 1985 · 1985
Earlier work this paper cites.
Autism and theory of mind in everyday life
Uta Frith. 1994 · 1994
Earlier work this paper cites.
Recognition of faux pas by normally developing children and children with asperger syndrome or high-functioning autism
Simon Baron-Cohen, Michelle O’riordan, Valerie Stone, Rosie Jones, and Kate Plaisted. 1999 · 1999
Earlier work this paper cites.
Conceptual alignment in conversation
Michael F Schober. 2005 · 2005
Earlier work this paper cites.
The effect of culture on perspective taking
Shali Wu and Boaz Keysar. 2007 · 2007
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme. 2013 · 2013
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
How can memory-augmented neural networks pass a false-belief task?
Erin Grant, Aida Nematzadeh, and Thomas L Griffiths. 2017 · 2017
Earlier work this paper cites.
Evaluating theory of mind in question answering
Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, and Tom Griffiths. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Modeling naive psychology of characters in simple commonsense stories
Hannah Rashkin, Antoine Bosselut, Maarten Sap, Kevin Knight, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
What makes reading comprehension questions easier?
Saku Sugawara, Kentaro Inui, Satoshi Sekine, and Akiko Aizawa. 2018 · 2018
Earlier work this paper cites.
Revisiting the evaluation of theory of mind through question answering
Matthew Le, Y-Lan Boureau, and Maximilian Nickel. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Experience grounds language
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020 · 2020
Cited alongside, same era.
Being deceived: Information asymmetry in second-order false belief tasks
Torben Braüner, Patrick Blackburn, and Irina Polyanskaya. 2020 · 2020
Cited alongside, same era.
Will I sound like me? improving persona consistency in dialogues through pragmatic self-consciousness
Hyunwoo Kim, Byeongchang Kim, and Gunhee Kim. 2020 · 2020
Cited alongside, same era.
What do theory-of-mind tasks actually measure? theory and practice
François Quesque and Yves Rossetti. 2020 · 2020
Cited alongside, same era.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Closest in time.
Ultrachat: A large-scale auto-generated multi-round dialogue data
Ning Ding, Yulin Chen, Bokai Xu, Shengding Hu, Yujia Qin, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023 · 2023
Closest in time.
Understanding social reasoning in language models with language models
Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah D Goodman. 2023 · 2023
Closest in time.
Zephyr 7b alpha
HuggingFace. 2023 · 2023
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021 · 2021
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022 · 2022
Cited alongside, same era.
Robots-dont-cry: Understanding falsely anthropomorphic utterances in dialog systems
David Gros, Yu Li, and Zhou Yu. 2022 · 2022
Cited alongside, same era.
Soda: Million-scale dialogue distillation with social commonsense contextualization
Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Le Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, and Yejin Choi. 2022 · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Closest in time.
Microsoft says new A.I. shows signs of human reasoning
Cade Metz. 2023 · 2023
Closest in time.
Boosting theory-of-mind performance in large language models via prompting
Shima Rahimi Moghaddam and Christopher J Honey. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Closest in time.
Minding language models’ (lack of) theory of mind: A plug-and-play multi-character belief tracker
Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov. 2023 · 2023
Closest in time.
How well do large language models perform on faux pas tests
Natalie Shapira, Guy Zwirn, and Yoav Goldberg. 2023b · 2023
Closest in time.
Ul2: Unifying language learning paradigms
Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, et al. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Large language models fail on trivial alterations to theory-of-mind tasks
Tomer Ullman. 2023 · 2023
Closest in time.
Can a machine know that we know what it knows?
Oliver Whang. 2023 · 2023
Closest in time.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023 · 2023
Closest in time.
How far are large language models from agents with theory-of-mind?
Pei Zhou, Aman Madaan, Srividya Pranavi Potharaju, Aditya Gupta, Kevin R McKee, Ari Holtzman, Jay Pujara, Xiang Ren, Swaroop Mishra, Aida Nematzadeh, et al. 2023 · 2023
Closest in time.