Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are trained to imitate humans to explain human decisions.
Mental models in cognitive science
Philip N Johnson-Laird. 1980 · 1980
Earlier work this paper cites.
Harvey Friedman’s research on the foundations of mathematics
Leo A Harrington, Michael D Morley, A Šcedrov, and Stephen G Simpson. 1985 · 1985
Earlier work this paper cites.
How people construct mental models
Allan Collins and Dedre Gentner. 1987 · 1987
Earlier work this paper cites.
Mental models as representations of discourse and text
Alan Garnham. 1987 · 1987
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Mental models, trust, and reliance: Exploring the effect of human perceptions on automation use
Andrea M Cassidy. 2009 · 2009
Earlier work this paper cites.
Mental models
Dedre Gentner and Albert L Stevens. 2014 · 2014
Earlier work this paper cites.
"why should i trust you?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
The promise and peril of human evaluation for model interpretability
Bernease Herman. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Do explanations make VQA models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh. 2018 · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. 2018 · 2018
Earlier work this paper cites.
Teaching categories to human learners with visual explanations
Oisin Mac Aodha, Shihan Su, Yuxin Chen, Pietro Perona, and Yisong Yue. 2018 · 2018
Earlier work this paper cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2018 · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Earlier work this paper cites.
Explainable ai for healthcare: from black box to interpretable models
Amina Adadi and Mohammed Berrada. 2020 · 2019
Earlier work this paper cites.
Beyond accuracy: The role of mental models in human-ai team performance
Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz. 2019 · 2019
Earlier work this paper cites.
The judicial demand for explainable artificial intelligence
Ashley Deeks. 2019 · 2019
Earlier work this paper cites.
Counterfactual visual explanations
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Earlier work this paper cites.
An evaluation of the human-interpretability of explanation
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. 2019 · 2019
Cited alongside, same era.
Faithful and customizable explanations of black box models
Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Jure Leskovec. 2019 · 2019
Cited alongside, same era.
Faithful multimodal explanation for visual question answering
Jialin Wu and Raymond Mooney. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2020
Cited alongside, same era.
Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
HILDIF: Interactive debugging of NLI models using influence functions
Hugo Zylberajch, Piyawat Lertvittayakumjorn, and Francesca Toni. 2021 · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Later among the works it cites.
Frame: Evaluating simulatability metrics for free-text rationales
Aaron Chan, Shaoliang Nie, Liang Tan, Xiaochang Peng, Hamed Firooz, Maziar Sanjabi, and Xiang Ren. 2022 · 2022
Later among the works it cites.
Rev: Information-theoretic evaluation of free-text rationales
Hanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji, Yejin Choi, and Swabha Swayamdipta. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
NILE : Natural language inference with faithful natural language explanations
Sawan Kumar and Partha Talukdar. 2020 · 2020
Cited alongside, same era.
Evaluating explanation methods for neural machine translation
Jierui Li, Lemao Liu, Huayang Li, Guanlin Li, Guoping Huang, and Shuming Shi. 2020 · 2020
Cited alongside, same era.
Wt5?! training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020 · 2020
Cited alongside, same era.
Antonia Creswell and Murray Shanahan. 2022 · 2022
Later among the works it cites.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022 · 2022
Later among the works it cites.
Explainable artificial intelligence for digital forensics
Stuart W Hall, Amin Sakzad, and Kim-Kwang Raymond Choo. 2022 · 2022
Later among the works it cites.
Large language models can self-improve
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022 · 2022
Later among the works it cites.
Towards faithful model explanation in nlp: A survey
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
Logical satisfiability of counterfactuals for faithful explanations in nli
Suzanna Sia, Anton Belyy, Amjad Almahairi, Madian Khabsa, Luke Zettlemoyer, and Lambert Mathias. 2022 · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Large language models are reasoners with self-verification
Yixuan Weng, Minjun Zhu, Shizhu He, Kang Liu, and Jun Zhao. 2022 · 2022
Later among the works it cites.
Interactive AI Model Debugging and Correction
Tongshuang Wu. 2022 · 2022
Later among the works it cites.
Can explanations be useful for calibrating black box models?
Xi Ye and Greg Durrett. 2022 · 2022
Later among the works it cites.
Interpreting language models with contrastive explanations
Kayo Yin and Graham Neubig. 2022 · 2022
Later among the works it cites.
Are machine rationales (not) useful to humans? measuring and improving human utility of free-text rationales
Brihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang, Yejin Choi, and Xiang Ren. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al. 2023 · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman. 2023 · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Closest in time.