Fetching the paper…
Reading the bibliography…
The escalating debate on AI's capabilities warrants developing reliable metrics to assess machine "intelligence".
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
An experimental study of apparent behavior
Fritz Heider and Marianne Simmel. 1944 · 1944
Earlier work this paper cites.
Knowledge and Belief: An Introduction to the Logic of the Two Notions
Jaakko Hintikka. 1962 · 1962
Earlier work this paper cites.
An argument for the identity theory
David K Lewis. 1966 · 1966
Earlier work this paper cites.
Logic and conversation
Herbert P Grice. 1975 · 1975
Earlier work this paper cites.
Computer power and human reason: From judgment to calculation
Joseph Weizenbaum. 1976 · 1976
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff. 1978 · 1978
Earlier work this paper cites.
Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception
Heinz Wimmer and Josef Perner. 1983 · 1983
Earlier work this paper cites.
Does the autistic child have a “theory of mind”?
Simon Baron-Cohen, Alan M Leslie, and Uta Frith. 1985 · 1985
Earlier work this paper cites.
Three-year-olds’ difficulty with false belief: The case for a conceptual deficit
Josef Perner, Susan R Leekam, and Heinz Wimmer. 1987 · 1987
Earlier work this paper cites.
Centering: A framework for modeling the local coherence of discourse
Barbara J. Grosz, Aravind K. Joshi, and Scott Weinstein. 1995 · 1995
Earlier work this paper cites.
Recognition of faux pas by normally developing children and children with asperger syndrome or high-functioning autism
Simon Baron-Cohen, Michelle O’riordan, Valerie Stone, Rosie Jones, and Kate Plaisted. 1999 · 1999
Earlier work this paper cites.
On seeing human: a three-factor theory of anthropomorphism
Nicholas Epley, Adam Waytz, and John T Cacioppo. 2007 · 2007
Earlier work this paper cites.
Language and theory of mind: Meta-analysis of the relation between language ability and false-belief understanding
Karen Milligan, Janet Wilde Astington, and Lisa Ain Dack. 2007 · 2007
Earlier work this paper cites.
The psychological meaning of words: Liwc and computerized text analysis methods
Yla R Tausczik and James W Pennebaker. 2010 · 2010
Earlier work this paper cites.
Anthropomorphism of computers: Is it mindful or mindless?
Youjeong Kim and S Shyam Sundar. 2012 · 2012
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme. 2013 · 2013
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merriënboer, Armand Joulin, and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
Commonsense interpretation of triangle behavior
Andrew Gordon. 2016 · 2016
Earlier work this paper cites.
A course in machine learning
Hal Daumé. 2017 · 2017
Earlier work this paper cites.
A formal theory of commonsense psychology: How people think people think
Andrew S Gordon and Jerry R Hobbs. 2017 · 2017
Earlier work this paper cites.
How can memory-augmented neural networks pass a false-belief task?
Erin Grant, Aida Nematzadeh, and Thomas L Griffiths. 2017 · 2017
Cited alongside, same era.
Detecting depression and mental illness on social media: an integrative review
Sharath Chandra Guntuku, David B Yaden, Margaret L Kern, Lyle H Ungar, and Johannes C Eichstaedt. 2017 · 2017
Cited alongside, same era.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Cited alongside, same era.
Evaluating theory of mind in question answering
Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, and Thomas L Griffiths. 2018 · 2018
Cited alongside, same era.
When choosing plausible alternatives, clever hans can be clever
Pride Kavumba, Naoya Inoue, Benjamin Heinzerling, Keshav Singh, Paul Reisert, and Kentaro Inui. 2019 · 2019
Cited alongside, same era.
Looking for the lighthouse: A systematic review of advanced theory-of-mind tests beyond preschool
Christopher Osterhaus and Sandra L Bosacki. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
The “problem” of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. 2022 · 2022
Later among the works it cites.
Neural theory-of-mind? on the limits of social intelligence in large lms
Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi. 2022 · 2022
Later among the works it cites.
“alexa, do you want to build a snowman?” characterizing playful requests to conversational agents
Chen Shani, Alexander Libov, Sofia Tolmach, Liane Lewin-Eytan, Yoelle Maarek, and Dafna Shahaf. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Revisiting the evaluation of theory of mind through question answering
Matthew Le, Y-Lan Boureau, and Maximilian Nickel. 2019 · 2019
Cited alongside, same era.
Socialiqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Do neural language models overcome reporting bias?
Vered Shwartz and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Unsupervised commonsense question answering with self-talk
Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021 · 2021
Cited alongside, same era.
Exploring roberta’s theory of mind through textual entailment
Michael Cohen. 2021 · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022 · 2022
Later among the works it cites.
Unifying language learning paradigms
Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023 · 2023
Closest in time.
Experimentology: An open science approach to experimental psychology methods
Michael C Frank, Mika Braginsky, Julie Cachia, Nicholas Coles, Tom Hardwicke, Robert Hawkins, Maya B Mathur, and Rondeline Williams. 2023 · 2023
Closest in time.
Theory of mind may have spontaneously emerged in large language models
Michal Kosinski. 2023 · 2023
Closest in time.
How not to test GPT-3
Gary Marcus and Ernest Davis. 2023 · 2023
Closest in time.
GPT-4 technical report
OpenAI. 2023 · 2023
Closest in time.
Closed ai models make bad baselines
Anna Rodgers. 2023 · 2023
Closest in time.
Evaluating humorous response generation to playful shopping requests
Natalie Shapira, Oren Kalinsky, Alex Libov, Chen Shani, and Sofia Tolmach. 2023a · 2023
Closest in time.
How well do large language models perform on faux pas tests
Natalie Shapira, Guy Zwirn, and Yoav Goldberg. 2023b · 2023
Closest in time.
Large language models fail on trivial alterations to theory-of-mind tasks
Tomer Ullman. 2023 · 2023
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023 · 2023
Closest in time.
Can a machine know that we know what it knows?
Oliver Whang. 2023 · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023 · 2023
Closest in time.