Fetching the paper…
Reading the bibliography…
Can we preserve the accuracy of neural models while also providing faithful explanations of model decisions to training data? We propose a "wrapper box'' pipeline: training a neural model as usual and then using its learned feature representation in classic, interpretable models to perform prediction.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 1910
Earlier work this paper cites.
Human problem solving
Allan Newel and Herbert A Simon. 1972 · 1972
Earlier work this paper cites.
Automatic indexing
Karen Spärck Jones. 1974 · 1974
Earlier work this paper cites.
Multidimensional binary search trees used for associative searching
Jon Louis Bentley. 1975 · 1975
Earlier work this paper cites.
An improved approximate formula for calculating sample sizes for comparing two binomial distributions
J. T. Casagrande, M. C. Pike, and P. G. Smith. 1978 · 1978
Earlier work this paper cites.
Principal components analysis (pca)
Andrzej Maćkiewicz and Waldemar Ratajczak. 1993 · 1993
Earlier work this paper cites.
Case-based reasoning: Foundational issues, methodological variations, and system approaches
Agnar Aamodt and Enric Plaza. 1994 · 1994
Earlier work this paper cites.
Case-based explanation of non-case-based learning methods
R Caruana, H Kangarloo, J D Dionisio, U Sinha, and D Johnson. 1999 · 1999
Earlier work this paper cites.
Capacity limits of information processing in the brain
René Marois and Jason Ivanoff. 2005 · 2005
Earlier work this paper cites.
Explanation in case-based Reasoning–Perspectives and goals
Frode Sørmo, Jörg Cassens, and Agnar Aamodt. 2005 · 2005
Earlier work this paper cites.
Twitter sentiment classification using distant supervision
Alec Go. 2009 · 2009
Earlier work this paper cites.
Explaining and improving model behavior with k nearest neighbor representations
Nazneen Fatema Rajani, Ben Krause, Wengpeng Yin, Tong Niu, Richard Socher, and Caiming Xiong. 2020 · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
The hewlett foundation: Automated essay scoring
Ben Hamner, Jaison Morgan, lynnvandev, Mark Shermis, and Tom Vander Ark. 2012 · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Case-Based Reasoning
Janet Kolodner. 2014 · 2014
Earlier work this paper cites.
Inside case-based explanation
Roger C Schank, Alex Kass, and Christopher K Riesbeck. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
"why should i trust you?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Case-based reasoning
Michael M Richter and Rosina O Weber. 2016 · 2016
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
k-nearest neighbor augmented neural networks for text classification
Zhiguo Wang, Wael Hamza, and Linfeng Song. 2017 · 2017
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (XAI)
Amina Adadi and Mohammed Berrada. 2018 · 2018
Earlier work this paper cites.
Interpretable machine learning in healthcare
M A Ahmad, C Eckert, and A Teredesai. 2018 · 2018
Earlier work this paper cites.
’it’s reducing a human being to a percentage’: Perceptions of justice in algorithmic decisions
Reuben Binns, Max Van Kleek, Michael Veale, Ulrik Lyngs, Jun Zhao, and Nigel Shadbolt. 2018 · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Hate speech dataset from a white supremacy forum
Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018 · 2018
Earlier work this paper cites.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Nicolas Papernot and Patrick McDaniel. 2018 · 2018
Earlier work this paper cites.
CARER: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. 2018 · 2018
Earlier work this paper cites.
Interpreting neural networks with nearest neighbors
Eric Wallace, Shi Feng, and Jordan Boyd-Graber. 2018 · 2018
Earlier work this paper cites.
Tem: Tree-enhanced embedding model for explainable recommendation
Xiang Wang, Xiangnan He, Fuli Feng, Liqiang Nie, and Tat-Seng Chua. 2018 · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Sik Kim, Ian E H Yen, and Pradeep Ravikumar. 2018 · 2018
Cited alongside, same era.
Efficient knn classification with different numbers of nearest neighbors
Shichao Zhang, Xuelong Li, Ming Zong, Xiaofeng Zhu, and Ruili Wang. 2018 · 2018
Cited alongside, same era.
The effects of example-based explanations in a machine learning interface
Carrie J Cai, Jonas Jongejan, and Jess Holbrook. 2019 · 2019
Cited alongside, same era.
Machine learning interpretability: A survey on methods and metrics
Diogo V. Carvalho, Eduardo M. Pereira, and Jaime S. Cardoso. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
Early-stopped neural networks are consistent
Ziwei Ji, Justin Li, and Matus Telgarsky. 2021 · 2021
Later among the works it cites.
What do we want from explainable artificial intelligence (XAI)? – a stakeholder perspective on XAI and a conceptual model guiding interdisciplinary XAI research
Markus Langer, Daniel Oster, Timo Speith, Holger Hermanns, Lena Kästner, Eva Schmidt, Andreas Sesing, and Kevin Baum. 2021 · 2021
Later among the works it cites.
Explanation-based human debugging of NLP models: A survey
Piyawat Lertvittayakumjorn and Francesca Toni. 2021 · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G Northcutt, Anish Athalye, and Jonas Mueller. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Explaining models: An empirical study of how explanations impact fairness judgment
Jonathan Dodge, Q. Vera Liao, Yunfeng Zhang, Rachel K. E. Bellamy, and Casey Dugan. 2019 · 2019
Cited alongside, same era.
Techniques for interpretable machine learning
Mengnan Du, Ninghao Liu, and Xia Hu. 2019 · 2019
Cited alongside, same era.
Interpretable image recognition with hierarchical prototypes
Peter Hase, Chaofan Chen, Oscar Li, and Cynthia Rudin. 2019 · 2019
Cited alongside, same era.
The right to explanation, explained
Margot E Kaminski. 2019 · 2019
Cited alongside, same era.
How case-based reasoning explains neural networks: A theoretical analysis of xai using post-hoc explanation-by-example from a survey of ann-cbr twin-systems
Mark T. Keane and Eoin M. Kenny. 2019 · 2019
Cited alongside, same era.
Interpreting black box predictions using fisher kernels
Rajiv Khanna, Been Kim, Joydeep Ghosh, and Sanmi Koyejo. 2019 · 2019
Cited alongside, same era.
SELFEXPLAIN: A self-explaining architecture for neural text classifiers
Dheeraj Rajagopal, Vidhisha Balachandran, Eduard H Hovy, and Yulia Tsvetkov. 2021 · 2021
Later among the works it cites.
The effects of explainability and causability on perception, trust, and acceptance: Implications for explainable AI
Donghee Shin. 2021 · 2021
Later among the works it cites.
Designing inherently interpretable machine learning models
Agus Sudjianto and Aijun Zhang. 2021 · 2021
Later among the works it cites.
Interactive label cleaning with example-based explanations
Stefano Teso, Andrea Bontempelli, Fausto Giunchiglia, and Andrea Passerini. 2021 · 2021
Later among the works it cites.
Are explanations helpful? a comparative study of the effects of explanations in AI-assisted decision-making
Xinru Wang and Ming Yin. 2021 · 2021
Later among the works it cites.
On sample based explanation methods for NLP: Faithfulness, efficiency and semantic evaluation
Wei Zhang, Ziming Huang, Yada Zhu, Guangnan Ye, Xiaodong Cui, and Fan Zhang. 2021 · 2021
Later among the works it cites.
Evaluating the quality of machine learning explanations: A survey on methods and metrics
Jianlong Zhou, Amir H. Gandomi, Fang Chen, and Andreas Holzinger. 2021 · 2021
Later among the works it cites.
HILDIF: Interactive debugging of NLI models using influence functions
Hugo Zylberajch, Piyawat Lertvittayakumjorn, and Francesca Toni. 2021 · 2021
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022 · 2022
Later among the works it cites.
ProtoTEx: Explaining model decisions with prototype tensors
Anubrata Das, Chitrank Gupta, Venelin Kovatchev, Matthew Lease, and Junyi Jessy Li. 2022 · 2022
Later among the works it cites.
Position: The case against case-based explanation
Jonathan Dodge. 2022 · 2022
Later among the works it cites.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Later among the works it cites.
Datamodels: Understanding predictions with data and data with predictions
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry. 2022 · 2022
Later among the works it cites.
A survey of algorithmic recourse: Contrastive explanations and consequential recommendations
Amir-Hossein Karimi, Gilles Barthe, Bernhard Schölkopf, and Isabel Valera. 2022 · 2022
Later among the works it cites.
Post-hoc interpretability for neural nlp: A survey
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2022 · 2022
Later among the works it cites.
SHAP-based explanation methods: A review for NLP interpretability
Edoardo Mosca, Ferenc Szigeti, Stella Tragianni, Daniel Gallagher, and Georg Groh. 2022 · 2022
Later among the works it cites.
Robust explainability: A tutorial on gradient-based attribution methods for deep neural networks
Ian E. Nielsen, Dimah Dera, Ghulam Rasool, Ravi P. Ramachandran, and Nidhal Carla Bouaynaya. 2022 · 2022
Later among the works it cites.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. 2022 · 2022
Later among the works it cites.
Intuitively assessing ml model reliability through example-based explanations and editing model inputs
Harini Suresh, Kathleen M Lewis, John Guttag, and Arvind Satyanarayan. 2022 · 2022
Later among the works it cites.
Understanding the role of human intuition on reliance in human-AI decision-making with explanations
Valerie Chen, Q Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal. 2023 · 2023
Closest in time.
IAEval: A comprehensive evaluation of instance attribution on natural language understanding
Peijian Gu, Yaozong Shen, Lijie Wang, Quan Wang, Hua Wu, and Zhendong Mao. 2023 · 2023
Closest in time.
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Closest in time.
Learning human-compatible representations for case-based decision support
Han Liu, Yizhou Tian, Chacha Chen, Shi Feng, Yuxin Chen, and Chenhao Tan. 2023 · 2023
Closest in time.
Human alignment of neural network representations
Lukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen, and Simon Kornblith. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Few-shot relation classification using clustering-based prototype modification
Mingtong Wen, Tingyu Xia, Bowen Liao, and Yuan Tian. 2023 · 2023
Closest in time.
How many and which training points would need to be removed to flip this prediction?
Jinghan Yang, Sarthak Jain, and Byron C Wallace. 2023 · 2023
Closest in time.
A critical survey on fairness benefits of explainable AI
Luca Deck, Jakob Schoeffer, Maria De-Arteaga, and Niklas Kühl. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Thomas Mesnard et al. 2024 · 2024
Closest in time.
Relabeling minimal training subset to flip a prediction
Jinghan Yang, Linjie Xu, and Lequan Yu. 2024 · 2024
Closest in time.