Fetching the paper…
Reading the bibliography…
In this paper we delve deep in the Transformer architecture by investigating two of its core components: self-attention and contextual embeddings.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 1905
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 1906
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 1908
Earlier work this paper cites.
On structural identifiability
R. Bellman and Karl Johan Åström · 1970
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F. Brown, Stephen A. Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer · 2003
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
Kristina Toutanova, Dan Klein, Christopher D. Manning, and Yoram Singer · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Extracting syntactic trees from transformer encoder self-attentions
David Marecek and Rudolf Rosa · 2018
Earlier work this paper cites.
Dissecting contextual word embeddings: Architecture and representation
Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih · 2018
Earlier work this paper cites.
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement
Nina Pörner, Hinrich Schütze, and Benjamin Roth · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann · 2018
Cited alongside, same era.
An analysis of attention mechanisms: The case of word sense disambiguation in neural machine translation
Gongbo Tang, Rico Sennrich, and Joakim Nivre · 2018
Cited alongside, same era.
Attending to mathematical language with transformers
Artit Wangperawong · 2018
Cited alongside, same era.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah · 2019
Closest in time.
Microsoft translator at wmt 2019: Towards large-scale document-level neural machine translation
Marcin Junczys-Dowmunt · 2019
Closest in time.
Attention is (not) all you need for commonsense reasoning
Tassilo Klein and Moin Nabi · 2019
Closest in time.
Open sesame: Getting inside bert’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank · 2019
Closest in time.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
Modeling localness for self-attention networks
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, and Tong Zhang · 2018
Cited alongside, same era.
A BERT Baseline for the Natural Questions
Chris Alberti, Kenton Lee, and Michael Collins · 2019
Cited alongside, same era.
Do transformer attention heads provide transparency in abstractive summarization?
Joris Baan, Maartje ter Hoeve, Marlies van der Wees, Anne Schuth, and Maarten de Rijke · 2019
Cited alongside, same era.
What does BERT look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of BERT
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda B. Viégas, and Martin Wattenberg · 2019
Cited alongside, same era.
Harshith Padigela, Hamed Zamani, and W. Bruce Croft · 2019
Closest in time.
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C Lipton · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Closest in time.
Is attention interpretable?
Sofia Serrano and Noah A. Smith · 2019
Closest in time.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Closest in time.
Visualizing attention in transformer-based language representation models
Jesse Vig · 2019
Closest in time.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Closest in time.
Adding interpretable attention to neural translation models improves word alignment
Thomas Zenkel, Joern Wuebker, and John DeNero · 2019
Closest in time.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Closest in time.