Fetching the paper…
Reading the bibliography…
Past work in natural language processing interpretability focused mainly on popular classification tasks while largely overlooking generation settings, partly due to a lack of dedicated tools.
Gradio: Hassle-free sharing and testing of ML models in the wild
Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, Abdulrahman Alfozan, and James Y. Zou. 2019 · 1906
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, T. J. Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Captum: A unified and generic model interpretability library for pytorch
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson. 2020 · 2009
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert Müller. 2010 · 2010
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013 · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Earlier work this paper cites.
"why should i trust you?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
A causal framework for explaining the predictions of black-box sequence-to-sequence models
David Alvarez-Melis and Tommi Jaakkola. 2017 · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Earlier work this paper cites.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Context-aware neural machine translation learns anaphora resolution
Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018 · 2018
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Earlier work this paper cites.
On measuring gender bias in translation of gender-neutral pronouns
Won Ik Cho, Ji Won Kim, Seok Min Kim, and Nam Soo Kim. 2019 · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
Saliency-driven word alignment interpretation for neural machine translation
Shuoyang Ding, Hainan Xu, and Philipp Koehn. 2019 · 2019
Earlier work this paper cites.
Towards understanding neural machine translation with word importance
Shilin He, Zhaopeng Tu, Xing Wang, Longyue Wang, Michael Lyu, and Shuming Shi. 2019 · 2019
Earlier work this paper cites.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema. 2020 · 2020
Cited alongside, same era.
A diagnostic study of explainability techniques for text classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 · 2020
Cited alongside, same era.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Cited alongside, same era.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Cited alongside, same era.
Is MAP decoding all you need? the inadequacy of the mode in neural machine translation
On the lack of robust interpretability of neural text classifiers
Muhammad Bilal Zafar, Michele Donini, Dylan Slack, Cedric Archambeau, Sanjiv Das, and Krishnaram Kenthapadi. 2021 · 2021
Later among the works it cites.
ferret: a framework for benchmarking explainers on transformers
Giuseppe Attanasio, Eliana Pastor, Chiara Di Bonaventura, and Debora Nozza. 2022 · 2022
Later among the works it cites.
“will you find these shortcuts?” a protocol for evaluating the faithfulness of input salience methods for text classification
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova. 2022 · 2022
Later among the works it cites.
Trustworthy social bias measurement
Rishi Bommasani and Percy Liang. 2022 · 2022
Later among the works it cites.
An empirical study on explanations in out-of-domain settings
George Chrysostomou and Nikolaos Aletras. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bryan Eikema and Wilker Aziz. 2020 · 2020
Cited alongside, same era.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
Inserting Information Bottlenecks for Attribution in Transformers
Zhiying Jiang, Raphael Tang, Ji Xin, and Jimmy Lin. 2020 · 2020
Cited alongside, same era.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020 · 2020
Cited alongside, same era.
Assessing gender bias in machine translation: a case study with google translate
Marcelo OR Prates, Pedro H Avelar, and Luís C Lamb. 2020 · 2020
Cited alongside, same era.
The language interpretability tool: Extensible, interactive visualizations and analysis for NLP models
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, and Ann Yuan. 2020 · 2020
Cited alongside, same era.
The tatoeba translation challenge – realistic data sets for low resource and multilingual MT
Jörg Tiedemann. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc Le, and Jason Wei. 2022 · 2022
Later among the works it cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Later among the works it cites.
David Dale, Elena Voita, Loïc Barrault, and Marta Ruiz Costa-jussà. 2022 · 2022
Later among the works it cites.
GPT3.int8(): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
How to measure gender bias in machine translation: Real-world oriented machine translators, multiple reference points
Anna Farkas and Renáta Németh. 2022 · 2022
Later among the works it cites.
Towards opening the black box of neural machine translation: Source and target interpretations of the transformer
Javier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano, and Marta R. Costa-jussà. 2022 · 2022
Later among the works it cites.
Influence functions for sequence tagging models
Sarthak Jain, Varun Manjunatha, Byron Wallace, and Ani Nenkova. 2022 · 2022
Later among the works it cites.
Analyzing the use of influence functions for instance-specific data filtering in neural machine translation
Tsz Kin Lam, Eva Hasler, and Felix Hieber. 2022 · 2022
Later among the works it cites.
Evaluating the faithfulness of importance measures in NLP by recursively masking allegedly important tokens and retraining
Andreas Madsen, Nicholas Meade, Vaibhav Adlakha, and Siva Reddy. 2022a · 2022
Later among the works it cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022 · 2022
Later among the works it cites.
GlobEnc: Quantifying global token attribution by incorporating the whole encoder layer in transformers
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, and Mohammad Taher Pilehvar. 2022 · 2022
Later among the works it cites.
You reap what you sow: On the challenges of bias evaluation under multilingual settings
Zeerak Talat, Aurélie Névéol, Stella Biderman, Miruna Clinciu, Manan Dey, Shayne Longpre, Sasha Luccioni, Maraim Masoud, Margaret Mitchell, Dragomir Radev, Shanya Sharma, Arjun Subramonian, Jaesung Tae, Samson Tan, Deepak Tunuguntla, and Oskar Van Der Wal. 2022 · 2022
Later among the works it cites.
Reducing hallucinations in neural machine translation with feature attribution
Joel Tang, M. Fomicheva, and Lucia Specia. 2022 · 2022
Later among the works it cites.
Undesirable biases in nlp: Averting a crisis of measurement
Oskar van der Wal, Dominik Bachmann, Alina Leidinger, Leendert van Maanen, Willem Zuidema, and Katrin Schulz. 2022 · 2022
Later among the works it cites.
Interpreting language models with contrastive explanations
Kayo Yin and Graham Neubig. 2022 · 2022
Later among the works it cites.
Explaining how transformers use context to build predictions
Javier Ferrando, Gerard I. G’allego, Ioannis Tsiamas, and Marta Ruiz Costa-jussà. 2023 · 2023
Closest in time.
Feed-forward blocks control contextualization in masked language models
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2023 · 2023
Closest in time.
Quantifying context mixing in transformers
Hosein Mohebbi, Willem H. Zuidema, Grzegorz Chrupała, and A. Alishahi. 2023 · 2023
Closest in time.
Attribution patching: Activation patching at industrial scale
Neel Nanda. 2023 · 2023
Closest in time.
Understanding and detecting hallucinations in neural machine translation via model introspection
Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, and Marine Carpuat. 2023 · 2023
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.