Fetching the paper…
Reading the bibliography…
In the last half-decade, the field of natural language processing (NLP) has undergone two major transitions: the switch to neural networks as the primary modeling paradigm and the homogenization of the training regime (pre-train, then fine-tune).
Tenney, Ian, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019 · 1905
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, Alex, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
On identifiability in transformers
Brunner, Gino, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2019 · 1908
Earlier work this paper cites.
Swayamdipta, Swabha, Matthew Peters, Brendan Roof, Chris Dyer, and Noah A Smith. 2019 · 1908
Earlier work this paper cites.
Do attention heads in bert track syntactic dependencies?
Htut, Phu Mon, Jason Phang, Shikha Bordia, and Samuel R Bowman. 2019 · 1911
Earlier work this paper cites.
Language
Bloomfield, Leonard. 1933 · 1933
Earlier work this paper cites.
Human Behavior and the Principle of Least Effort: An Introduction to Human Ecology
Zipf, G.K. 1949 · 1949
Earlier work this paper cites.
Syntactic Structures
Chomsky, Noam. 1957 · 1957
Earlier work this paper cites.
Éléments de syntaxe structurale
Tesnière, Lucien. 1959 · 1959
Earlier work this paper cites.
Aspects of the Theory of Syntax
Chomsky, Noam. 1965 · 1965
Earlier work this paper cites.
Lectures on Government and Binding
Chomsky, Noam. 1981 · 1981
Earlier work this paper cites.
Dependency Syntax: Theory and Practice
Mel’čuk, Igor. 1988 · 1988
Earlier work this paper cites.
The greenbergian word order correlations
Dryer, Matthew S. 1992 · 1992
Earlier work this paper cites.
The Minimalist Program
Chomsky, Noam. 1995 · 1995
Earlier work this paper cites.
Functionalism and grammar
Givón, Talmy. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Linguistic complexity: Locality of syntactic dependencies
Gibson, Edward. 1998 · 1998
Earlier work this paper cites.
The dependency locality theory: A distance-based theory of linguistic complexity
Gibson, Edward et al. 2000 · 2000
Earlier work this paper cites.
Efficiency and complexity in grammars
Hawkins, John A. 2004 · 2004
Earlier work this paper cites.
A systematic assessment of syntactic generalization in neural language models
Hu, Jennifer, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger P Levy. 2020 · 2005
Earlier work this paper cites.
Generating typed dependency parses from phrase structure parses
de Marneffe, Marie-Catherine, Bill MacCartney, and Christopher D. Manning. 2006 · 2006
Earlier work this paper cites.
The myth of language universals: Language diversity and its importance for cognitive science
Evans, Nicholas and Stephen C Levinson. 2009 · 2009
Earlier work this paper cites.
Unbounded dependency recovery for parser evaluation
Rimell, Laura, Stephen Clark, and Mark Steedman. 2009 · 2009
Earlier work this paper cites.
The cultural origins of human cognition
Tomasello, Michael. 2009 · 2009
Earlier work this paper cites.
Measuring association between labels and free-text rationales
Wiegreffe, Sarah, Ana Marasović, and Noah A Smith. 2020 · 2010
Earlier work this paper cites.
On language ‘utility’: Processing complexity and communicative efficiency
Jaeger, T Florian and Harry Tily. 2011 · 2011
Earlier work this paper cites.
Who did what to whom? a contrastive study of syntacto-semantic dependencies
Ivanova, Angelina, Stephan Oepen, Lilja Øvrelid, and Dan Flickinger. 2012 · 2012
Earlier work this paper cites.
Pham, Thang M, Trung Bui, Long Mai, and Anh Nguyen. 2020 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, Karen, Andrea Vedaldi, and Andrew Zisserman. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, Jeffrey, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Trends in syntactic parsing: Anticipation, bayesian estimation, and good-enough parsing
Traxler, Matthew J. 2014 · 2014
Cited alongside, same era.
Recurrent neural network grammars
Dyer, Chris, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
" why should i trust you?" explaining the predictions of any classifier
Ribeiro, Marco Tulio, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
The argument from the poverty of the stimulus
Lasnik, Howard and Jeffrey L Lidz. 2017 · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Lundberg, Scott M and Su-In Lee. 2017 · 2017
Cited alongside, same era.
Attention is not Explanation
Jain, Sarthak and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
Linguistic knowledge and transferability of contextual representations
Liu, Nelson F., Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019a · 2019
Later among the works it cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, Tom, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Later among the works it cites.
Explain yourself! leveraging language models for commonsense reasoning
Rajani, Nazneen Fatema, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Later among the works it cites.
Is attention interpretable?
Serrano, Sofia and Noah A. Smith. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Axiomatic attribution for deep networks
Sundararajan, Mukund, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
The anthropological setting of polysynthesis
Trudgill, Peter. 2017 · 2017
Cited alongside, same era.
e-snli: Natural language inference with natural language explanations
Camburu, Oana-Maria, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Cited alongside, same era.
Sud or surface-syntactic universal dependencies: An annotation scheme near-isomorphic to ud
Gerdes, Kim, Bruno Guillaume, Sylvain Kahane, and Guy Perrier. 2018 · 2018
Cited alongside, same era.
Colorless green recurrent networks dream hierarchically
Gulordava, Kristina, Piotr Bojanowski, Édouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Gururangan, Suchin, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Tenney, Ian, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Attention is not not explanation
Wiegreffe, Sarah and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Quantifying attention flow in transformers
Abnar, Samira and Willem Zuidema. 2020 · 2020
Later among the works it cites.
A diagnostic study of explainability techniques for text classification
Atanasova, Pepa, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 · 2020
Later among the works it cites.
Dependency locality as an explanatory principle for word order
Futrell, Richard, Roger P Levy, and Edward Gibson. 2020 · 2020
Later among the works it cites.
Syntaxgym: An online platform for targeted evaluation of language models
Gauthier, Jon, Jennifer Hu, Ethan Wilcox, Peng Qian, and Roger Levy. 2020 · 2020
Later among the works it cites.
Universals of word order reflect optimization of grammars for efficient communication
Hahn, Michael, Dan Jurafsky, and Richard Futrell. 2020 · 2020
Later among the works it cites.
Do neural language models show preferences for syntactic formalisms?
Kulmizev, Artur, Vinit Ravishankar, Mostafa Abdou, and Joakim Nivre. 2020 · 2020
Later among the works it cites.
Syntactic structure distillation pretraining for bidirectional encoders
Kuncoro, Adhiguna, Lingpeng Kong, Daniel Fried, Dani Yogatama, Laura Rimell, Chris Dyer, and Phil Blunsom. 2020 · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Manning, Christopher D, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020 · 2020
Later among the works it cites.
A tale of a probe and a parser
Maudslay, Rowan Hall, Josef Valvoda, Tiago Pimentel, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Composition is the core driver of the language-selective network
Mollica, Francis, Matthew Siegelman, Evgeniia Diachek, Steven T Piantadosi, Zachary Mineroff, Richard Futrell, Hope Kean, Peng Qian, and Evelina Fedorenko. 2020 · 2020
Later among the works it cites.
Pareto probing: Trading off accuracy for complexity
Pimentel, Tiago, Naomi Saphra, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Sinha, Koustuv, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. 2020 · 2020
Later among the works it cites.
Blimp: The benchmark of linguistic minimal pairs for english
Warstadt, Alex, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2020 · 2020
Later among the works it cites.
On the proper role of linguistically-oriented deep net analysis in linguistic theorizing
Baroni, Marco. 2021 · 2021
Closest in time.
Probing classifiers: Promises, shortcomings, and alternatives
Belinkov, Yonatan. 2021 · 2021
Closest in time.
Demystifying neural language models’ insensitivity to word-order
Clouatre, Louis, Prasanna Parthasarathi, Amal Zouaq, and Sarath Chandar. 2021 · 2021
Closest in time.
Bert & family eat word salad: Experiments with text understanding
Gupta, Ashim, Giorgi Kvernadze, and Vivek Srikumar. 2021 · 2021
Closest in time.
Aligning faithful interpretations with their social attribution
Jacovi, Alon and Yoav Goldberg. 2021 · 2021
Closest in time.
Syntactic structure from deep learning
Linzen, Tal and Marco Baroni. 2021 · 2021
Closest in time.
Universal dependencies
de Marneffe, Marie, Christopher D. Manning, Joakim Nivre, and Daniel Zeman. 2021 · 2021
Closest in time.
Refining targeted syntactic evaluation of language models
Newman, Benjamin, Kai-Siang Ang, Julia Gong, and John Hewitt. 2021 · 2021
Closest in time.
Sinha, Koustuv, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021 · 2021
Closest in time.
Infusing finetuning with semantic dependencies
Wu, Zhaofeng, Hao Peng, and Noah A Smith. 2021 · 2021
Closest in time.