Fetching the paper…
Reading the bibliography…
Explainability methods for NLP systems encounter a version of the fundamental problem of causal inference: for a given ground-truth input text, we never truly observe the counterfactual texts necessary for isolating the causal effects of model representations on outputs.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 1910
Earlier work this paper cites.
Statistics and causal inference
Paul W Holland · 1986
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Counterfactual explanations for machine learning: A review
Sahil Verma, John Dickerson, and Keegan Hines · 2010
Earlier work this paper cites.
API design for machine learning software: Experiences from the scikit-learn project
Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake VanderPlas, Arnaud Joly, Brian Holt, and Gaël Varoquaux · 2013
Earlier work this paper cites.
Safe and interpretable machine learning: A methodological review
Clemens Otte · 2013
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Striving for simplicity: the all convolutional net
Jost Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al · 2015
Earlier work this paper cites.
Interactive and Interpretable Machine Learning Models for Human Machine Collaboration
Been Kim · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Layer-wise relevance propagation for neural networks with local renormalization layers
Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, and Wojciech Samek · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro · 2016
Earlier work this paper cites.
"Why Should I Trust You?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
European Union regulations on algorithmic decision-making and a “right to explanation”
Bryce Goodman and Seth Flaxman · 2017
Earlier work this paper cites.
Inherent trade-offs in the fair determination of risk scores
Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan · 2017
Earlier work this paper cites.
Learning Important Features through Propagating Activation Differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Cited alongside, same era.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres · 2018
Cited alongside, same era.
The mythos of model interpretability
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy · 2020
Later among the works it cites.
Interpretable Machine Learning
Christoph Molnar · 2020
Later among the works it cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Later among the works it cites.
Discovering the compositional structure of vector representations with role learning networks
Paul Soulos, R. Thomas McCoy, Tal Linzen, and Paul Smolensky · 2020
Later among the works it cites.
Investigating gender bias in language models using causal mediation analysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zachary C. Lipton · 2018
Cited alongside, same era.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Counterfactual visual explanations
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
AI-mediated communication: How the perception that profile text was written by AI affects trustworthiness
Maurice Jakesch, Megan French, Xiao Ma, Jeffrey T Hancock, and Mor Naaman · 2019
Cited alongside, same era.
Metalearners for estimating heterogeneous treatment effects using machine learning
Sören R Künzel, Jasjeet S Sekhon, Peter J Bickel, and Bin Yu · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber · 2020
Later among the works it cites.
On completeness-aware concept-based explanations in deep neural networks
Chih-Kuan Yeh, Been Kim, Sercan Arik, Chun-Liang Li, Tomas Pfister, and Pradeep Ravikumar · 2020
Later among the works it cites.
Sparse interventions in language models with differentiable masking
Nicola De Cao, Leon Schmid, Dieuwke Hupkes, and Ivan Titov · 2021
Later among the works it cites.
Expanding explainability: Towards social transparency in AI systems
Upol Ehsan, Q Vera Liao, Michael Muller, Mark O Riedl, and Justin D Weisz · 2021
Later among the works it cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2021
Later among the works it cites.
CausaLM: Causal model explanation through counterfactual language models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart · 2021
Later among the works it cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Later among the works it cites.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld · 2021
Later among the works it cites.
CEBaB: Estimating the causal effects of real-world concepts on NLP model behavior
Eldar David Abraham, Karel D’Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu · 2022
Closest in time.
Testing pre-trained language models’ understanding of distributivity via causal mediation analysis
Pangbo Ban, Yifan Jiang, Tianran Liu, and Shane Steinert-Threlkeld · 2022
Closest in time.
Inducing causal structure for interpretable neural networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah Goodman, and Christopher Potts · 2022
Closest in time.
Unit testing for concepts in neural networks
Charles Lovering and Ellie Pavlick · 2022
Closest in time.