Fetching the paper…
Reading the bibliography…
Large pre-trained neural networks such as BERT have had great recent success in NLP, motivating a growing body of research investigating what aspects of language they are able to learn from unlabeled data.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Sarthak Jain and Byron C. Wallace. 2019 · 1902
Earlier work this paper cites.
Visualizing attention in transformer-based language models
Jesse Vig. 2019 · 1904
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 1905
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 1905
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 1905
Earlier work this paper cites.
Multidimensional scaling by optimizing goodness of fit to a nonmetric hypothesis
Joseph B Kruskal. 1964 · 1964
Earlier work this paper cites.
Building a large annotated corpus of english: The Penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
Stanford’s multi-pass sieve coreference resolution system at the conll-2011 shared task
Heeyoung Lee, Yves Peirsman, Angel Chang, Nathanael Chambers, Mihai Surdeanu, and Dan Jurafsky. 2011 · 2011
Earlier work this paper cites.
Conll-2012 shared task: Modeling multilingual unrestricted coreference in ontonotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012 · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le. 2015 · 2015
Earlier work this paper cites.
Learning anaphoricity and antecedent ranking features for coreference resolution
Sam Joshua Wiseman, Alexander Matthew Rush, Stuart Merrill Shieber, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Tree-to-sequence attentional neural machine translation
Akiko Eriguchi, Kazuma Hashimoto, and Yoshimasa Tsuruoka. 2016 · 2016
Cited alongside, same era.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Does string-based neural mt learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Cited alongside, same era.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Cited alongside, same era.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James R. Glass. 2017 · 2017
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Daniel Jurafsky. 2018 · 2018
Later among the works it cites.
Extracting syntactic trees from transformer encoder self-attentions
David Marecek and Rudolf Rosa. 2018 · 2018
Later among the works it cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Deep rnns encode soft hierarchical syntax
Terra Blevins, Omer Levy, and Luke S. Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Exploiting attention to reveal shortcomings in memory models
Kaylee Burns, Aida Nematzadeh, Alison Gopnik, and Thomas L. Griffiths. 2018 · 2018
Cited alongside, same era.
Syntax-directed attention for neural machine translation
Kehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita, and Tiejun Zhao. 2018 · 2018
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, Germán Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Linguistically-informed self-attention for semantic role labeling
Emma Strubell, Patrick Verga, Daniel Andor, David I Weiss, and Andrew McCallum. 2018 · 2018
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2018 · 2018
Later among the works it cites.
Multi-head attention with disagreement regularization
Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Later among the works it cites.
Context-aware neural machine translation learns anaphora resolution
Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018 · 2018
Later among the works it cites.
Language modeling teaches you more syntax than translation does: Lessons learned through auxiliary task analysis
Kelly W. Zhang and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Finding syntax with structural probes
John Hewitt and Christopher D. Manning. 2019 · 2019
Closest in time.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, M. Peters, and Noah A. Smith. 2019 · 2019
Closest in time.