Fetching the paper…
Reading the bibliography…
A possible explanation for the impressive performance of masked language model (MLM) pre-training is that such models have learned to represent the syntactic structures prevalent in classical NLP pipelines.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Distributional structure
Zellig S Harris. 1954 · 1954
Earlier work this paper cites.
Syntactic structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
Some universals of grammar with particular reference to the order of meaningful elements
Joseph Greenberg. 1963 · 1963
Earlier work this paper cites.
Universal coding, information, prediction, and estimation
Jorma Rissanen. 1984 · 1984
Earlier work this paper cites.
Indexing by latent semantic analysis
Scott Deerwester, Susan T Dumais, George W Furnas, Thomas K Landauer, and Richard Harshman. 1990 · 1990
Earlier work this paper cites.
The Greenbergian word order correlations
Matthew S Dryer. 1992 · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A solution to plato’s problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge
Thomas K Landauer and Susan T Dumais. 1997 · 1997
Earlier work this paper cites.
Adverbs and functional heads: A cross-linguistic perspective
Guglielmo Cinque. 1999 · 1999
Earlier work this paper cites.
Training batchnorm and only batchnorm: On the expressive power of random features in cnns
Jonathan Frankle, David J Schwab, and Ari S Morcos. 2020 · 2003
Earlier work this paper cites.
Catching the drift: Probabilistic content models, with applications to generation and summarization
Regina Barzilay and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
R Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006 · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B Dolan. 2007 · 2007
Earlier work this paper cites.
Modeling local coherence: An entity-based approach
Regina Barzilay and Mirella Lapata. 2008 · 2008
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009 · 2009
Earlier work this paper cites.
English web treebank
Ann Bies, Justin Mott, Colin Warner, and Seth Kulick. 2012 · 2012
Earlier work this paper cites.
Thang M. Pham, Trung Bui, Long Mai, and Anh Nguyen. 2020 · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Cutting recursive autoencoder trees
Christian Scheible and Hinrich Schütze. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2013
Earlier work this paper cites.
A gold standard dependency corpus for english
Natalia Silveira, Timothy Dozat, Marie-Catherine De Marneffe, Samuel R Bowman, Miriam Connor, John Bauer, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Improving distributional similarity with lessons learned from word embeddings
Omer Levy, Yoav Goldberg, and Ido Dagan. 2015 · 2015
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D Manning. 2017 · 2017
Earlier work this paper cites.
SentEval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela. 2018 · 2018
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Transforming question answering datasets into natural language inference datasets
Dorottya Demszky, Kelvin Guu, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018a · 2018
Cited alongside, same era.
PAWS: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He. 2019 · 2019
Later among the works it cites.
Learning to few-shot learn across diverse natural language classification tasks
Trapit Bansal, Rishikesh Jha, and Andrew McCallum. 2020 · 2020
Later among the works it cites.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Later among the works it cites.
A tale of a probe and a parser
Rowan Hall Maudslay, Josef Valvoda, Tiago Pimentel, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
spaCy: Industrial-strength Natural Language Processing in Python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Later among the works it cites.
OCNLI: Original Chinese Natural Language Inference
Hai Hu, Kyle Richardson, Liang Xu, Lu Li, Sandra Kübler, and Lawrence Moss. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Do language models understand anything? on the ability of LSTMs to understand negative polarity items
Jaap Jumelet and Dieuwke Hupkes. 2018 · 2018
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
Performance impact caused by hidden bias of training data for recognizing textual entailment
Masatoshi Tsuchiya. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Pre-training without Natural Images
Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, and Yutaka Satoh. 2020 · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020 · 2020
Later among the works it cites.
Learning Music Helps You Read: Using transfer to study linguistic structure in language models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
Pareto probing: Trading off accuracy for complexity
Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell. 2020a · 2020
Later among the works it cites.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Masked Language Model Scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020 · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Later among the works it cites.
BLiMP: A benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020a · 2020
Later among the works it cites.
Learning which features matter: RoBERTa acquires a preference for linguistic generalizations (eventually)
Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, and Samuel R. Bowman. 2020b · 2020
Later among the works it cites.
Efficient second-order TreeCRF for neural dependency parsing
Yu Zhang, Zhenghua Li, and Min Zhang. 2020 · 2020
Later among the works it cites.
Syntactic perturbations reveal representational correlates of hierarchical phrase structure in pretrained language models
Matteo Alleman, Jonathan Mamou, Miguel A Del Rio, Hanlin Tang, Yoon Kim, and SueYeon Chung. 2021 · 2021
Closest in time.
Probing classifiers: Promises, shortcomings, and alternatives
Yonatan Belinkov. 2021 · 2021
Closest in time.
Is supervised syntactic parsing beneficial for language understanding tasks? an empirical investigation
Goran Glavaš and Ivan Vulić. 2021 · 2021
Closest in time.
Bert & family eat word salad: Experiments with text understanding
Ashim Gupta, Giorgi Kvernadze, and Vivek Srikumar. 2021 · 2021
Closest in time.
Language models use monotonicity to assess NPI licensing
Jaap Jumelet, Milica Denic, Jakub Szymanik, Dieuwke Hupkes, and Shane Steinert-Threlkeld. 2021 · 2021
Closest in time.
Can transformer models measure coherence in text: Re-thinking the shuffle test
Philippe Laban, Luke Dai, Lucas Bandarkar, and Marti A. Hearst. 2021 · 2021
Closest in time.
Mechanisms for handling nested dependencies in neural-network language models and humans
Yair Lakretz, Dieuwke Hupkes, Alessandra Vergallito, Marco Marelli, Marco Baroni, and Stanislas Dehaene. 2021 · 2021
Closest in time.
What context features can transformer language models use?
Joe O’Connor and Jacob Andreas. 2021 · 2021
Closest in time.
Deep subjecthood: Higher-order grammatical features in multilingual BERT
Isabel Papadimitriou, Ethan A. Chi, Richard Futrell, and Kyle Mahowald. 2021 · 2021
Closest in time.
Sometimes we want ungrammatical translations
Prasanna Parthasarathi, Koustuv Sinha, Joelle Pineau, and Adina Williams. 2021 · 2021
Closest in time.
Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021 · 2021
Closest in time.
Reservoir transformers
Sheng Shen, Alexei Baevski, Ari Morcos, Kurt Keutzer, Michael Auli, and Douwe Kiela. 2021 · 2021
Closest in time.
UnNatural Language Inference
Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. 2021 · 2021
Closest in time.
On the interplay between fine-tuning and composition in transformers
Lang Yu and Allyson Ettinger. 2021 · 2021
Closest in time.
Revisiting few-sample BERT fine-tuning
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger, and Yoav Artzi. 2021 · 2021
Closest in time.