Fetching the paper…
Reading the bibliography…
A hallmark property of explainable AI models is the ability to teach other agents, communicating knowledge of how to perform a task.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff · 1978
Earlier work this paper cites.
Theory of mind for a humanoid robot
Brian Scassellati · 2002
Earlier work this paper cites.
Theory of mind for learning and teaching: The nature and role of explanation
Henry M Wellman and Kristin H Lagattuta · 2004
Earlier work this paper cites.
The explanation game: Towards prediction explainability through sparse communication
Marcos Treviso and André FT Martins · 2004
Earlier work this paper cites.
Evaluating explainable ai: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal · 2005
Earlier work this paper cites.
Constructing a language: A usage-based theory of language acquisition
Michael Tomasello · 2005
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld · 2006
Earlier work this paper cites.
Joint mind modeling for explanation generation in complex human-robot collaborative tasks
Xiaofeng Gao, Ran Gong, Yizhou Zhao, Shu Wang, Tianmin Shu, and Song-Chun Zhu · 2007
Earlier work this paper cites.
Are visual explanations useful? a case study in model-in-the-loop prediction
Eric Chu, Deb Roy, and Jacob Andreas · 2007
Earlier work this paper cites.
Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal · 2010
Earlier work this paper cites.
Evaluating explanations: How much do explanations from the teacher aid students?
Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C Lipton, Graham Neubig, and William W Cohen · 2012
Earlier work this paper cites.
Theory of mind
Stephanie M Carlson, Melissa A Koenig, and Madeline B Harms · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Reasoning about pragmatics with neural listeners and speakers
Jacob Andreas and Dan Klein · 2016
Earlier work this paper cites.
An implemented theory of mind to improve human-robot shared plans execution
Sandra Devin and Rachid Alami · 2016
Earlier work this paper cites.
Pragmatic language interpretation as probabilistic inference
Noah D Goodman and Michael C Frank · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Cited alongside, same era.
Do explanations make VQA models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh · 2018
Cited alongside, same era.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton · 2018
Cited alongside, same era.
Evaluating theory of mind in question answering
Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, and Tom Griffiths · 2018
Cited alongside, same era.
Knife: Knowledge distillation with free-text rationales
Aaron Chan, Zhiyuan Zeng, Wyatt Lake, Brihi Joshi, Hanjie Chen, and Xiang Ren · 2022
Later among the works it cites.
Reframing human-ai collaboration for generating free-text explanations
Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi · 2022
Later among the works it cites.
Are hard examples also harder to explain? a study with human and model-generated explanations
Swarnadeep Saha, Peter Hase, Nazneen Rajani, and Mohit Bansal · 2022
Later among the works it cites.
PaLM: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Cited alongside, same era.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Cited alongside, same era.
Young children consider the expected utility of others’ learning to decide what to teach
Sophie Bridgers, Julian Jara-Ettinger, and Hyowon Gweon · 2020
Cited alongside, same era.
Few-shot language coordination by modeling theory of mind
Hao Zhu, Graham Neubig, and Yonatan Bisk · 2021
Cited alongside, same era.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Cited alongside, same era.
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Later among the works it cites.
Theory of mind may have spontaneously emerged in large language models
Michal Kosinski · 2023
Closest in time.
Large language models fail on trivial alterations to theory-of-mind tasks
Tomer Ullman · 2023
Closest in time.
Clever hans or neural theory of mind? stress testing social reasoning in large language models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz · 2023
Closest in time.
Specializing smaller language models towards multi-step reasoning
Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, and Tushar Khot · 2023
Closest in time.
Post hoc explanations of language models can improve language models
Jiaqi Ma, Dylan Slack, Asma Ghandeharioun, Sameer Singh, Himabindu Lakkaraju, et al · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Closest in time.
Brihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang, Yejin Choi, and Xiang Ren · 2023
Closest in time.
User adaptive language learning chatbots with a curriculum
Kun Qian, Ryan Shea, Yu Li, Luke Kutszik Fryer, and Zhou Yu · 2023
Closest in time.
Computational language acquisition with theory of mind
Andy Liu, Hao Zhu, Emmy Liu, Yonatan Bisk, and Graham Neubig · 2023
Closest in time.
Boosting theory-of-mind performance in large language models via prompting
Shima Rahimi Moghaddam and Christopher J Honey · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.