Fetching the paper…
Reading the bibliography…
The self-rationalising capabilities of large language models (LLMs) have been explored in restricted settings, using task/specific data sets.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Brian W Matthews. 1975 · 1975
Earlier work this paper cites.
Self-explanations: How students study and use examples in learning to solve problems
Michelene TH Chi, Miriam Bassok, Matthew W Lewis, Peter Reimann, and Robert Glaser. 1989 · 1989
Earlier work this paper cites.
Conversational processes and causal explanation
Denis J Hilton. 1990 · 1990
Earlier work this paper cites.
Use of prior beliefs in the assignment of causal roles: Causal powers versus regularity-based accounts
Peter A White. 1995 · 1995
Earlier work this paper cites.
Measuring individual differences in implicit cognition: the implicit association test
Anthony G Greenwald, Debbie E McGhee, and Jordan LK Schwartz. 1998 · 1998
Earlier work this paper cites.
The Shadows and Shallows of Explanation
Robert A. Wilson and Frank Keil. 1998 · 1998
Earlier work this paper cites.
The misunderstood limits of folk science: An illusion of explanatory depth
Leonid Rozenblit and Frank Keil. 2002 · 2002
Earlier work this paper cites.
What lies beneath? Understanding the limits of understanding. Thinking and Seeing: Visual Metacognition in Adults and Children
F Keil, L Rozenblit, and C Mills. 2004 · 2004
Earlier work this paper cites.
Explanation and understanding
Frank C Keil. 2006 · 2006
Earlier work this paper cites.
The structure and function of explanations
Tania Lombrozo. 2006 · 2006
Earlier work this paper cites.
LIREx: Augmenting Language Inference with Relevant Explanation
Xinyan Zhao and V. G. Vinod Vydiswaran. 2020 · 2012
Earlier work this paper cites.
How People Explain Action (and Autonomous Intelligent Systems Should Too)
Maartje M. A. de Graaf and Bertram F. Malle. 2017 · 2017
Earlier work this paper cites.
A Roadmap for a Rigorous Science of Interpretability
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
The promise and peril of human evaluation for model interpretability
Bernease Herman. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI)
Amina Adadi and Mohammed Berrada. 2018 · 2018
Earlier work this paper cites.
e-SNLI: Natural Language Inference with Natural Language Explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Hallucinations in neural machine translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2018 · 2018
Earlier work this paper cites.
Menaka Narayanan, Emily Chen, Jeffrey He, Been Kim, Sam Gershman, and Finale Doshi-Velez. 2018 · 2018
Earlier work this paper cites.
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2018 · 2018
Cited alongside, same era.
Improving language understanding with unsupervised learning
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Automated Rationale Generation: A Technique for Explainable AI and Its Effects on Human Perceptions
Upol Ehsan, Pradyumna Tambwekar, Larry Chan, Brent Harrison, and Mark O. Riedl. 2019 · 2019
Cited alongside, same era.
Towards explainable NLP: A generative explanation framework for text classification
Hui Liu, Qingyu Yin, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2019 · 2019
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Later among the works it cites.
Few-shot self-rationalization with natural language prompts
Ana Marasovic, Iz Beltagy, Doug Downey, and Matthew Peters. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
On the diversity and limits of human explanations
Chenhao Tan. 2022 · 2022
Later among the works it cites.
Label Studio: Data labeling software
Maxim Tkachenko, Mikhail Malyuk, Andrey Holmanyuk, and Nikolai Liubimov. 2020-2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brent Mittelstadt, Chris Russell, and Sandra Wachter. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Cited alongside, same era.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
WT5?! Training Text-to-Text Models to Explain their Predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2020
Cited alongside, same era.
Explanations for CommonsenseQA: New Dataset and Models
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021 · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. 2021 · 2021
Cited alongside, same era.
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022 · 2022
Later among the works it cites.
Finetuned Language Models are Zero-Shot Learners
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a · 2022
Later among the works it cites.
Reframing human-AI collaboration for generating free-text explanations
Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi. 2022 · 2022
Later among the works it cites.
The unreliability of explanations in few-shot prompting for textual reasoning
Xi Ye and Greg Durrett. 2022 · 2022
Later among the works it cites.
Factuality challenges in the era of large language models
Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha, Tanmoy Chakraborty, Giovanni Luca Ciampaglia, David Corney, Renee DiResta, Emilio Ferrara, Scott Hale, Alon Halevy, Eduard Hovy, Heng Ji, Filippo Menczer, Ruben Miguez, Preslav Nakov, Dietram Scheufele, Shivam Sharma, and Giovanni Zagni. 2023 · 2023
Later among the works it cites.
Good Explanations in Explainable Artificial Intelligence (XAI): Evidence from Human Explanatory Reasoning
Ruth M.J. Byrne. 2023 · 2023
Later among the works it cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Later among the works it cites.
R Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L Griffiths. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023 · 2023
Later among the works it cites.
Wolfgang Stammer, Felix Friedrich, David Steinmann, Hikaru Shindo, and Kristian Kersting. 2023 · 2023
Later among the works it cites.
Stanford Alpaca: An Instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023 · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Later among the works it cites.