Fetching the paper…
Reading the bibliography…
Intelligent agents such as robots are increasingly deployed in real-world, safety-critical settings.
The need for user models in generating expert system explanations
Robert Kass, Tim Finin, et al · 1988
Earlier work this paper cites.
Agents that learn to explain themselves
W Lewis Johnson · 1994
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Natural language generation enhances human decision-making with uncertain information
Dimitra Gkatzia, Oliver Lemon, and Verena Rieser · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
An exploratory study on the benefits of using natural language for explaining fuzzy rule-based systems
Jose M Alonso, Alejandro Ramos-Soto, Ehud Reiter, and Kees van Deemter · 2017
Earlier work this paper cites.
Interpretability via model extraction
Osbert Bastani, Carolyn Kim, and Hamsa Bastani · 2017
Earlier work this paper cites.
Survey of machine learning algorithms for disease diagnostic
Meherwar Fatima and Maruf Pasha · 2017
Earlier work this paper cites.
Distilling a neural network into a soft decision tree
Nicholas Frosst and Geoffrey Hinton · 2017
Earlier work this paper cites.
Improving robot controller transparency through autonomous policy explanation
Bradley Hayes and Julie A Shah · 2017
Earlier work this paper cites.
Hierarchical and interpretable skill acquisition in multi-task reinforcement learning
Tianmin Shu, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Verifiable reinforcement learning via policy extraction
Osbert Bastani, Yewen Pu, and Armando Solar-Lezama · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom · 2018
Earlier work this paper cites.
Programmatically interpretable reinforcement learning
Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, and Swarat Chaudhuri · 2018
Earlier work this paper cites.
Summarizing agent strategies
Ofra Amir, Finale Doshi-Velez, and David Sarne · 2019
Earlier work this paper cites.
Probabilistic multimodal modeling for human-robot interaction tasks
Joseph Campbell, Simon Stepputtis, and Heni Ben Amor · 2019
Cited alongside, same era.
Techniques for interpretable machine learning
Mengnan Du, Ninghao Liu, and Xia Hu · 2019
Cited alongside, same era.
Automated rationale generation: a technique for explainable ai and its effects on human perceptions
Upol Ehsan, Pradyumna Tambwekar, Larry Chan, Brent Harrison, and Mark O Riedl · 2019
Cited alongside, same era.
Xai—explainable artificial intelligence
David Gunning, Mark Stefik, Jaesik Choi, Timothy Miller, Simone Stumpf, and Guang-Zhong Yang · 2019
Cited alongside, same era.
Generating justifications for norm-related agent decisions
Daniel Kasenberg, Antonio Roque, Ravenna Thielstrom, Meia Chita-Tegmark, and Matthias Scheutz · 2019
Cited alongside, same era.
Interactive explanations: Diagnosis and repair of reinforcement learning based agent behaviors
Christian Arzate Cruz and Takeo Igarashi · 2021
Later among the works it cites.
Evaluating artificial social intelligence in an urban search and rescue task environment
Jared T Freeman, Lixiao Huang, Matt Woods, and Stephen J Cauffman · 2021
Later among the works it cites.
Edge: Explaining deep reinforcement learning policies
Wenbo Guo, Xian Wu, Usmann Khan, and Xinyu Xing · 2021
Later among the works it cites.
Few-shot self-rationalization with natural language prompts
Ana Marasović, Iz Beltagy, Doug Downey, and Matthew E Peters · 2021
Later among the works it cites.
To what extent do human explanations of model behavior align with actual model behavior?
Grusha Prasad, Yixin Nie, Mohit Bansal, Robin Jia, Douwe Kiela, and Adina Williams · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Lewis, Katia Sycara, and Illah Nourbakhsh · 2019
Cited alongside, same era.
Toward interpretable deep reinforcement learning with linear model u-trees
Guiliang Liu, Oliver Schulte, Wang Zhu, and Qingcan Li · 2019
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Cited alongside, same era.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher · 2019
Cited alongside, same era.
Verbal explanations for deep reinforcement learning neural networks with attention on extracted features
Xinzhi Wang, Shengcheng Yuan, Hui Zhang, Michael Lewis, and Katia Sycara · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Natural language rationales with full-stack visual reasoning: From pixels to semantic frames to commonsense graphs
Ana Marasović, Chandra Bhagavatula, Jae sung Park, Ronan Le Bras, Noah A. Smith, and Yejin Choi · 2020
Cited alongside, same era.
Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi · 2021
Later among the works it cites.
Explanations from large language models make small reasoners better
Shiyang Li, Jianshu Chen, Yelong Shen, Zhiyu Chen, Xinlu Zhang, Zekun Li, Hong Wang, Jing Qian, Baolin Peng, Yi Mao, et al · 2022
Later among the works it cites.
Knowledge-grounded self-rationalization via extractive and natural language explanations
Bodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, and Julian Mcauley · 2022
Later among the works it cites.
Why? why not? when? visual explanations of agent behaviour in reinforcement learning
Aditi Mishra, Utkarsh Soni, Jinbin Huang, and Chris Bryan · 2022
Later among the works it cites.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Closest in time.
Explainable action advising for multi-agent reinforcement learning
Yue Guo, Joseph Campbell, Simon Stepputtis, Ruiyu Li, Dana Hughes, Fei Fang, and Katia Sycara · 2023
Closest in time.
Thought cloning: Learning to think while acting by imitating human thinking
Shengran Hu and Jeff Clune · 2023
Closest in time.
A novel policy-graph approach with natural language and counterfactual abstractions for explaining reinforcement learning agents
Tongtong Liu, Joe McCalmon, Thai Le, Md Asifur Rahman, Dongwon Lee, and Sarra Alqahtani · 2023
Closest in time.
Sources of hallucination by large language models on inference tasks
Nick McKenna, Tianyi Li, Liang Cheng, Mohammad Javad Hosseini, Mark Johnson, and Mark Steedman · 2023
Closest in time.
Simon Stepputtis, Joseph Campbell, Yaqi Xie, Zhengyang Qi, Wenxin Sharon Zhang, Ruiyi Wang, Sanketh Rangreji, Michael Lewis, and Katia Sycara · 2023
Closest in time.
Concept learning for interpretable multi-agent reinforcement learning
Renos Zabounidis, Joseph Campbell, Simon Stepputtis, Dana Hughes, and Katia P Sycara · 2023
Closest in time.