Fetching the paper…
Reading the bibliography…
Attention is a key component of Transformers, which have recently achieved considerable success in natural language processing.
Assessing BERT’s Syntactic Abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Adding Interpretable Attention to Neural Translation Models Improves Word Alignment
Thomas Zenkel, Joern Wuebker, and John DeNero. 2019 · 1901
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Attention Interpretability Across NLP Tasks
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019 · 1909
Earlier work this paper cites.
Do Attention Heads in BERT Track Syntactic Dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R. Bowman. 2019 · 1911
Earlier work this paper cites.
Improved Statistical Alignment Models
Franz Josef Och and Hermann Ney. 2000 · 2000
Earlier work this paper cites.
A Primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2002
Earlier work this paper cites.
A Systematic Comparison of Various Statistical Alignment Models
Franz Josef Och and Hermann Ney. 2003 · 2003
Earlier work this paper cites.
Telling BERT’s full story: from Local Attention to Global Aggregation
Damian Pascual, Gino Brunner, and Roger Wattenhofer. 2020 · 2004
Earlier work this paper cites.
AER: Do we need to \CJK@punctchar
David Vilar, Maja Popović, and Hermann Ney. 2006 · 2006
Earlier work this paper cites.
A Simple, Fast, and Effective Reparameterization of IBM Model 2
Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013 · 2013
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
An Analysis of Encoder Representations in Transformer-Based Machine Translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Cited alongside, same era.
An Analysis of Attention Mechanisms: The Case of Word Sense Disambiguation in Neural Machine Translation
Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2018 · 2018
Cited alongside, same era.
What Does BERT Look At? An Analysis of BERT’s Attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
From Balustrades to Pierre Vinken: Looking for Syntax in Transformer Self-Attentions
David Mareček and Rudolf Rosa. 2019 · 2019
Later among the works it cites.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Visualizing and Measuring the Geometry of BERT
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019 · 2019
Later among the works it cites.
Is Attention Interpretable?
Sofia Serrano and Noah A Smith. 2019 · 2019
Later among the works it cites.
What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Saliency-driven Word Alignment Interpretation for Neural Machine Translation
Shuoyang Ding, Hainan Xu, and Philipp Koehn. 2019 · 2019
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C Wallace. 2019 · 2019
Cited alongside, same era.
What Does BERT Learn about the Structure of Language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Cited alongside, same era.
Revealing the Dark Secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
On the Word Alignment from Neural Machine Translation
Xintong Li, Guanlin Li, Lemao Liu, Max Meng, and Shuming Shi. 2019 · 2019
Cited alongside, same era.
Open Sesame: Getting Inside BERT’s Linguistic Knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Do NLP Models Know Numbers? Probing Numeracy in Embeddings
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019 · 2019
Later among the works it cites.
Attention is not not Explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
On Identifiability in Transformers
Gino Brunner, Yang Liu, Damián Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Closest in time.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.
Learning to Deceive with Attention-Based Explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C Lipton. 2020 · 2020
Closest in time.