Fetching the paper…
Reading the bibliography…
Why do models often attend to salient words, and how does this evolve throughout training? We approximate model training as a two stage process: early on in training when the attention weights are uniform, the model learns to translate individual input word `i` to `o` if they co-occur frequently.
Sofia Serrano and Noah A Smith. 2019b · 1906
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019b · 1908
Earlier work this paper cites.
Fine-grained sentiment analysis with faithful attention
Ruiqi Zhong, Steven Shao, and Kathleen McKeown. 2019 · 1908
Earlier work this paper cites.
Attention interpretability across nlp tasks
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019 · 1909
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, and Robert L Mercer. 1993 · 1993
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. 2015 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Cited alongside, same era.
Multi30K: Multilingual English-German image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Cited alongside, same era.
Analogs of linguistic structure in deep representations
Jacob Andreas and Dan Klein. 2017 · 2017
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee. 2018 · 2018
Cited alongside, same era.
HotFlip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler. 2018 · 2018
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Şimşekli, Levent Sagun, and Mert Gurbuzbalaban. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Cited alongside, same era.
S eq 2s eq-v is: A visual debugging tool for sequence-to-sequence models
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch, Adam Perer, Hanspeter Pfister, and Alexander M Rush. 2018 · 2018
Cited alongside, same era.
On identifiability in transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2019 · 2019
Cited alongside, same era.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019a
Cited in the paper.
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019a · 2019
Later among the works it cites.
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C. Lipton. 2020 · 2020
Later among the works it cites.
Understanding attention for text classification
Xiaobing Sun and Wei Lu. 2020 · 2020
Later among the works it cites.