What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Later among the works it cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Later among the works it cites.
Open sesame: Getting inside BERT’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Later among the works it cites.
From balustrades to pierre vinken: Looking for syntax in transformer self-attentions
David Mareček and Rudolf Rosa. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
compare-mt: A tool for holistic comparison of language generation systems
Graham Neubig, Zi-Yi Dou, Junjie Hu, Paul Michel, Danish Pruthi, and Xinyi Wang. 2019 · 2019
Later among the works it cites.
Promoting the knowledge of source syntax in transformer nmt is not needed
Thuong Hai Pham, Dominik Macháček, and Ondřej Bojar. 2019 · 2019
Later among the works it cites.
Challenge test sets for MT evaluation
Maja Popović and Sheila Castilho. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
The MuCoW test suite at WMT 2019: Automatically harvested multilingual contrastive word sense disambiguation test sets for machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2019 · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
Quantity doesn’t buy quality syntax with neural language models
Marten van Schijndel, Aaron Mueller, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Revisiting low-resource neural machine translation: A case study
Rico Sennrich and Biao Zhang. 2019 · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Later among the works it cites.
Encoders help you disambiguate word senses in neural machine translation
Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2019 · 2019
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Later among the works it cites.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Elena Voita, Rico Sennrich, and Ivan Titov. 2019a · 2019
Later among the works it cites.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli. 2019 · 2019
Later among the works it cites.
Leveraging local and global patterns for self-attention networks
Mingzhou Xu, Derek F. Wong, Baosong Yang, Yue Zhang, and Lidia S. Chao. 2019 · 2019
Later among the works it cites.
Convolutional self-attention networks
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, and Zhaopeng Tu. 2019 · 2019
Later among the works it cites.
On identifiability in transformers
Gino Brunner, Yang Liu, Damián Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Closest in time.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Closest in time.
Hard-coded gaussian attention for neural machine translation
Original
Weiqiu You, Simeng Sun, and Mohit Iyyer. 2020 · 2020
Closest in time.