Fetching the paper…
Reading the bibliography…
Much of explainable AI research treats explanations as a means for model inspection.
A fuzzy relative of the isodata process and its use in detecting compact well-separated clusters
Joseph C Dunn · 1973
Earlier work this paper cites.
Well-separated clusters and optimal fuzzy partitions
Joseph C Dunn · 1974
Earlier work this paper cites.
Categorization and representation of physics problems by experts and novices
Michelene TH Chi, Paul J Feltovich, and Robert Glaser · 1981
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Yann LeCun, Bernhard Boser, John Denker, Donnie Henderson, Richard Howard, Wayne Hubbard, and Lawrence Jackel · 1989
Earlier work this paper cites.
Eliciting self-explanations improves understanding
Michelene TH Chi, Nicholas De Leeuw, Mei-Hung Chiu, and Christian LaVancher · 1994
Earlier work this paper cites.
No free lunch theorems for optimization
David H Wolpert and William G Macready · 1997
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Making things happen: A theory of causal explanation
James Woodward · 2005
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
The influence of reflective self-explanations on problem-solving performance
Kyungbin Kwon and David H Jonassen · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Self-Reflecting Methods of Learning Research , pp. 3011–3015
Michaela Gläser-Zikuda · 2012
Earlier work this paper cites.
Comparative effects of test-enhanced learning and self-explanation on long-term retention
Douglas P Larsen, Andrew C Butler, and Henry L Roediger III · 2013
Earlier work this paper cites.
Systematic reflection: Implications for learning from failures and successes
Shmuel Ellis, Bernd Carette, Frederik Anseel, and Filip Lievens · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
VQA: visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Self-explanation, an instructional strategy to foster clinical reasoning in medical students
Martine Chamberland and Sílvia Mamede · 2015
Earlier work this paper cites.
Interpretation of prediction models using the input gradient
Yotam Hechtlinger · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi S. Jaakkola · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Measuring machine intelligence through visual question answering
C. Lawrence Zitnick, Aishwarya Agrawal, Stanislaw Antol, Margaret Mitchell, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
End-to-end differentiable proving
Tim Rocktäschel and Sebastian Riedel · 2017
Earlier work this paper cites.
Right for the right reasons: Training differentiable models by constraining their explanations
Andrew Slavin Ross, Michael C. Hughes, and Finale Doshi-Velez · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases
Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M. Summers · 2017
Earlier work this paper cites.
Explicit reasoning over end-to-end neural architectures for visual question answering
Somak Aditya, Yezhou Yang, and Chitta Baral · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
David Alvarez-Melis and Tommi S. Jaakkola · 2018
Earlier work this paper cites.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2018
Earlier work this paper cites.
Theories on self-reflection in education
Anna Belobrovy · 2018
Earlier work this paper cites.
Inducing self-explanation: A meta-analysis
Kiran Bisra, Qing Liu, John C Nesbit, Farimah Salimi, and Philip H Winne · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom · 2018
Earlier work this paper cites.
Learning from examples via self-explanations
Michelene TH Chi · 2018
Earlier work this paper cites.
Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions
Oscar Li, Hao Liu, Chaofan Chen, and Cynthia Rudin · 2018
Earlier work this paper cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach · 2018
Earlier work this paper cites.
Interpretable basis decomposition for visual explanation
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2018
Earlier work this paper cites.
No free lunch theorem: A review
Stavros P Adam, Stamatios-Aggelos N Alexandropoulos, Panos M Pardalos, and Michael N Vrahatis · 2019
Cited alongside, same era.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov · 2019
Cited alongside, same era.
Machine Learning Interpretability: A Survey on Methods and Metrics
Diogo V Carvalho, Eduardo M Pereira, and Jaime S Cardoso · 2019
Cited alongside, same era.
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Y. Zou, and Been Kim · 2019
Cited alongside, same era.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Explanatory learning: Beyond empiricism in neural networks
Antonio Norelli, Giorgio Mariani, Luca Moschella, Andrea Santilli, Giambattista Parascandolo, Simone Melzi, and Emanuele Rodolà · 2022
Later among the works it cites.
Evaluating explanations: How much do explanations from the teacher aid students?
Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C. Lipton, Graham Neubig, and William W. Cohen · 2022
Later among the works it cites.
Explainable deep learning: A field guide for the uninitiated
Gabrielle Ras, Ning Xie, Marcel van Gerven, and Derek Doran · 2022
Later among the works it cites.
Explainability via causal self-talk
Nicholas A. Roy, Junkyung Kim, and Neil C. Rabinowitz · 2022
Later among the works it cites.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
Learning not to learn: Training deep neural networks with biased data
Byungju Kim, Hyunwoo Kim, Kyungsu Kim, Sungjin Kim, and Junmo Kim · 2019
Cited alongside, same era.
Enriching visual with verbal explanations for relational concepts–combining lime with aleph
Johannes Rabold, Hannah Deininger, Michael Siebers, and Ute Schmid · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin · 2019
Cited alongside, same era.
Taking a HINT: leveraging explanations to make vision and language models more grounded
Ramprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Shalini Ghosh, Larry P. Heck, Dhruv Batra, and Devi Parikh · 2019
Cited alongside, same era.
Explanatory interactive machine learning
Stefano Teso and Kristian Kersting · 2019
Cited alongside, same era.
Learning to repair: Repairing model output errors after deployment using a dynamic memory of feedback
Niket Tandon, Aman Madaan, Peter Clark, and Yiming Yang · 2022
Later among the works it cites.
Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks
Lin Wang and Kuk-Jin Yoon · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Later among the works it cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Later among the works it cites.
Badr AlKhamissi, Siddharth Verma, Ping Yu, Zhijing Jin, Asli Celikyilmaz, and Mona T. Diab · 2023
Closest in time.
ILLUME: rationalizing vision-language models through human interactions
Manuel Brack, Patrick Schramowski, Björn Deiseroth, and Kristian Kersting · 2023
Closest in time.
The role of causality in explainable artificial intelligence
Gianluca Carloni, Andrea Berti, and Sara Colantonio · 2023
Closest in time.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch · 2023
Closest in time.
In search of verifiability: Explanations rarely enable complementary performance in ai-advised decision making
Raymond Fok and Daniel S Weld · 2023
Closest in time.
Thought cloning: Learning to think while acting by imitating human thinking
Shengran Hu and Jeff Clune · 2023
Closest in time.
Curious replay for model-based adaptation
Isaac Kauvar, Chris Doyle, Linqi Zhou, and Nick Haber · 2023
Closest in time.
Post hoc explanations of language models can improve language models
Satyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun, Sameer Singh, and Himabindu Lakkaraju · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Closest in time.
DERA: enhancing large language model completions with dialog-enabled resolving agents
Varun Nair, Elliot Schumacher, Geoffrey J. Tso, and Anitha Kannan · 2023
Closest in time.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang · 2023
Closest in time.
Concept-based explainable artificial intelligence: A survey
Eleonora Poeta, Gabriele Ciravegna, Eliana Pastor, Tania Cerquitelli, and Elena Baralis · 2023
Closest in time.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Closest in time.
Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities
Waddah Saeed and Christian W. Omlin · 2023
Closest in time.
Reflective-net: Learning from explanations
Johannes Schneider and Michalis Vlachos · 2023
Closest in time.
α \alpha ilp: thinking visual scenes as differentiable logic programs
Hikaru Shindo, Viktor Pfanschilling, Devendra Singh Dhami, and Kristian Kersting · 2023
Closest in time.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David D. Cox, Yiming Yang, and Chuang Gan · 2023
Closest in time.
Leveraging explanations in interactive machine learning: An overview
Stefano Teso, Öznur Alkan, Wolfgang Stammer, and Elizabeth Daly · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
Medmnist v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification
Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Closest in time.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang · 2023
Closest in time.
Interpretable concept bottlenecks to align reinforcement learning agents
Quentin Delfosse, Sebastian Sztwiertnia, Wolfgang Stammer, Mark Rothermel, and Kristian Kersting · 2024
Closest in time.
Going beyond XAI: A systematic survey for explanation-guided learning
Yuyang Gao, Siyi Gu, Junji Jiang, Sungsoo Ray Hong, Dazhou Yu, and Liang Zhao · 2024
Closest in time.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang · 2024
Closest in time.
Generative judge for evaluating alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu · 2024
Closest in time.
REFINER: reasoning feedback on intermediate representations
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings · 2024
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston · 2024
Closest in time.