Fetching the paper…
Reading the bibliography…
Pretrained language models have achieved a new state of the art on many NLP tasks, but there are still many open questions about how and why they work so well.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
No training required: Exploring random encoders for sentence classification
John Wieting and Douwe Kiela. 2019 · 1901
Earlier work this paper cites.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019b · 1905
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 1905
Earlier work this paper cites.
Open sesame: Getting inside bert’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 1906
Earlier work this paper cites.
Sensebert: Driving some sense into bert
Yoav Levine, Barak Lenz, Or Dagan, Dan Padnos, Or Sharir, Shai Shalev-Shwartz, Amnon Shashua, and Yoav Shoham. 2019 · 1908
Earlier work this paper cites.
An experiment in computational discrimination of english word senses
Ezra Black. 1988 · 1988
Earlier work this paper cites.
One sense per discourse
William A Gale, Kenneth W Church, and David Yarowsky. 1992 · 1992
Earlier work this paper cites.
Word-sense disambiguation using statistical models of Roget’s categories trained on large corpora
David Yarowsky. 1992 · 1992
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Semantic classes and syntactic ambiguity
Philip Resnik. 1993 · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
Christiane Fellbaum. 1998 · 1998
Earlier work this paper cites.
Integrating subject field codes into WordNet
Bernardo Magnini and Gabriela Cavaglià. 2000 · 2000
Earlier work this paper cites.
SENSEVAL-2: Overview
Philip Edmonds and Scott Cotton. 2001 · 2001
Earlier work this paper cites.
Discriminative training methods for hidden Markov models: Theory and experiments with perceptron algorithms
Michael Collins. 2002 · 2002
Earlier work this paper cites.
Supersense tagging of unknown nouns in wordnet
Massimiliano Ciaramita and Mark Johnson. 2003 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
The English all-words task
Benjamin Snyder and Martha Palmer. 2004 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Learning semantic classes for word sense disambiguation
Upali Sathyajith Kohomban and Wee Sun Lee. 2005 · 2005
Earlier work this paper cites.
Exploring the automatic selection of basic level concepts
Rubén Izquierdo Beviá, Armando Suárez Cueto, and Germán Rigau Claramunt. 2007 · 2007
Earlier work this paper cites.
An empirical study on class-based word sense disambiguation
Rubén Izquierdo, Armando Suárez, and German Rigau. 2009 · 2009
Cited alongside, same era.
Speech and Language Processing (2Nd Edition)
Daniel Jurafsky and James H. Martin. 2009 · 2009
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Glove: Global Vectors for Word Representation
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Later among the works it cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Kawin Ethayarajh. 2019 · 2019
Later among the works it cites.
Visualizing and understanding the effectiveness of BERT
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2019 · 2019
Later among the works it cites.
Designing and interpreting probes with control tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Semi-supervised sequence learning
Andrew M. Dai and Quoc V. Le. 2015 · 2015
Cited alongside, same era.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Does string-based neural MT learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks
Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2017 · 2017
Cited alongside, same era.
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Later among the works it cites.
What would elsa do? freezing layers during transformer fine-tuning
Jaejun Lee, Raphael Tang, and Jimmy Lin. 2019 · 2019
Later among the works it cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Later among the works it cites.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Cross-lingual alignment of contextual word embeddings, with applications to zero-shot dependency parsing
Tal Schuster, Ori Ram, Regina Barzilay, and Amir Globerson. 2019 · 2019
Later among the works it cites.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. 2019 · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Later among the works it cites.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Elena Voita, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 2019
Later among the works it cites.
Probing for semantic classes: Diagnosing the meaning content of word embeddings
Yadollah Yaghoobzadeh, Katharina Kann, T. J. Hazen, Eneko Agirre, and Hinrich Schütze. 2019 · 2019
Later among the works it cites.
On identifiability in transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Closest in time.
Establishing Strong Baselines for the New Decade: Sequence Tagging, Syntactic and Semantic Parsing with BERT
Han He and Jinho D. Choi. 2020 · 2020
Closest in time.
What happens to bert embeddings during fine-tuning?
Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020 · 2020
Closest in time.
Rare words: A major problem for contextualized embeddings and how to fix it by attentive mimicking
Timo Schick and Hinrich Schütze. 2020 · 2020
Closest in time.