Fetching the paper…
Reading the bibliography…
We seek to understand how the representations of individual tokens and the structure of the learned feature space evolve between layers in deep neural networks under different learning objectives.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Relations between two sets of variates
Harold Hotelling. 1936 · 1936
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek. 1999 · 1999
Earlier work this paper cites.
Estimation of entropy and mutual information
Liam Paninski. 2003 · 2003
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014 · 2014
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015 · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Does string-based neural mt learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Understanding and improving morphological learning in the neural machine translation decoder
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, and Stephan Vogel. 2017 · 2017
Earlier work this paper cites.
The representational geometry of word meanings acquired by neural machine translation models
Felix Hill, Kyunghyun Cho, Sébastien Jean, and Y Bengio. 2017 · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017 · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
A* CCG parsing with a supertag and dependency factored model
Masashi Yoshikawa, Hiroshi Noji, and Yuji Matsumoto. 2017 · 2017
Cited alongside, same era.
The lazy encoder: A fine-grained analysis of the role of morphology in neural machine translation
Arianna Bisazza and Clara Tump. 2018 · 2018
Cited alongside, same era.
The IWSLT 2018 Evaluation Campaign
Jan Niehues, Ronaldo Cattoni, Sebastian Stüker, Mauro Cettolo, Marco Turchi, and Marcello Federico. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Later among the works it cites.
Assessing generative models via precision and recall
Mehdi S. M. Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. 2018 · 2018
Later among the works it cites.
Language modeling teaches you more than translation does: Lessons learned through auxiliary syntactic task analysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep RNNs encode soft hierarchical syntax
Terra Blevins, Omer Levy, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Findings of the 2018 conference on machine translation (wmt18)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Insights on representational similarity in neural networks with canonical correlation
Ari S. Morcos, Maithra Raghu, and Samy Bengio. 2018 · 2018
Cited alongside, same era.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017a
Cited in the paper.
Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks
Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2017b
Cited in the paper.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019a
Cited in the paper.
Kelly Zhang and Samuel Bowman. 2018 · 2018
Later among the works it cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019 · 2019
Closest in time.
Understanding learning dynamics of language models with SVCCA
Naomi Saphra and Adam Lopez. 2019 · 2019
Closest in time.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Closest in time.