Fetching the paper…
Reading the bibliography…
Modern machine learning models are opaque, and as a result there is a burgeoning academic subfield on methods that explain these models' behavior.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2019
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 1909
Earlier work this paper cites.
Evolutionary principles in self-referential learning. on learning now to learn: The meta-meta-meta…-hook
Jurgen Schmidhuber · 1987
Earlier work this paper cites.
Trueskill(tm): A bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Scaffolding in teacher–student interaction: A decade of research
Janneke Van de Pol, Monique Volman, and Jos Beishuizen · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Evaluating explanations: How much do explanations from the teacher aid students?
Danish Pruthi, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C. Lipton, Graham Neubig, and William W. Cohen · 2012
Earlier work this paper cites.
Extraction of salient sentences from labelled documents
Misha Denil, Alban Demiraj, and Nando de Freitas · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek · 2015
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola · 2016
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C. Lipton · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Explaining recurrent neural network predictions in sentiment analysis
Leila Arras, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Do explanations make VQA models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh · 2018
Earlier work this paper cites.
Learning to explain: An information-theoretic perspective on model interpretation
Jianbo Chen, Le Song, Martin Wainwright, and Michael Jordan · 2018
Cited alongside, same era.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Flax: A neural network library and ecosystem for JAX, 2020
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Later among the works it cites.
Interpretation of NLP models through input marginalization
Siwon Kim, Jihun Yi, Eunji Kim, and Sungroh Yoon · 2020
Later among the works it cites.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui · 2020
Later among the works it cites.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saliency-driven word alignment interpretation for neural machine translation
Shuoyang Ding, Hainan Xu, and Philipp Koehn · 2019
Cited alongside, same era.
Generalized inner loop meta-learning
Edward Grefenstette, Brandon Amos, Denis Yarats, Phu Mon Htut, Artem Molchanov, Franziska Meier, Douwe Kiela, Kyunghyun Cho, and Soumith Chintala · 2019
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace · 2019
Cited alongside, same era.
Human-grounded evaluations of explanation methods for text classification
Piyawat Lertvittayakumjorn and Francesca Toni · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Cited alongside, same era.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André F. T. Martins · 2019
Cited alongside, same era.
Marcos Treviso and André F. T. Martins · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova · 2021
Later among the works it cites.
Should we be pre-training? an argument for end-task aware training as an alternative
Lucio M. Dery, Paul Michel, Ameet S. Talwalkar, and Graham Neubig · 2021
Later among the works it cites.
The Eval4NLP shared task on explainable quality estimation: Overview and results
Marina Fomicheva, Piyawat Lertvittayakumjorn, Wei Zhao, Steffen Eger, and Yang Gao · 2021
Later among the works it cites.
SPECTRA: Sparse structured text rationalization
Nuno M. Guerreiro and André F. T. Martins · 2021
Later among the works it cites.
Aligning Faithful Interpretations with their Social Attribution
Alon Jacovi and Yoav Goldberg · 2021
Later among the works it cites.
Order in the court: Explainable ai methods prone to disagreement
Michael Neely, Stefan F Schouten, Maurits JR Bleeker, and Ana Lucic · 2021
Later among the works it cites.
Teaching with commentaries
Aniruddh Raghu, Maithra Raghu, Simon Kornblith, David Duvenaud, and Geoffrey Hinton · 2021
Later among the works it cites.
Imagenet-21k pretraining for the masses
T. Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor · 2021
Later among the works it cites.
IST-unbabel 2021 submission for the explainable quality estimation shared task
Marcos Treviso, Nuno M. Guerreiro, Ricardo Rei, and André F. T. Martins · 2021
Later among the works it cites.
Rationales for sequential predictions
Keyon Vafa, Yuntian Deng, David Blei, and Alexander Rush · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld · 2021
Later among the works it cites.
Meta learning for knowledge distillation
Wangchunshu Zhou, Canwen Xu, and Julian McAuley · 2021
Later among the works it cites.
Explain, edit, and understand: Rethinking user study design for evaluating model explanations
Siddhant Arora, Danish Pruthi, Norman Sadeh, William Cohen, Zachary Lipton, and Graham Neubig · 2022
Closest in time.
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt · 2022
Closest in time.