Fetching the paper…
Reading the bibliography…
One reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. 2019 · 1906
Earlier work this paper cites.
Mutual exclusivity as a challenge for neural networks
Kanishk Gandhi and Brenden M Lake. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
The formation of learning sets
Harry F Harlow. 1949 · 1949
Earlier work this paper cites.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Brian W. Matthews. 1975 · 1975
Earlier work this paper cites.
Lectures on government and binding
Noam Chomsky. 1981 · 1981
Earlier work this paper cites.
Quantifying inductive bias: Ai learning algorithms and valiant’s learning framework
David Haussler. 1988 · 1988
Earlier work this paper cites.
Syntactic Theory: A Formal Introduction , 2 edition
Ivan A. Sag, Thomas Wasow, and Emily M. Bender. 2003 · 2003
Earlier work this paper cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2003
Earlier work this paper cites.
Which one is the dax? achieving mutual exclusivity with neural networks
Kristina Gulordava, Thomas Brochhagen, and Gemma Boleda. 2020 · 2004
Earlier work this paper cites.
When does data augmentation help generalization in nlp?
Rohan Jha, Charles Lovering, and Ellie Pavlick. 2020 · 2004
Earlier work this paper cites.
Syntactic data augmentation increases robustness to inference heuristics
Junghyun Min, R Thomas McCoy, Dipanjan Das, Emily Pitler, and Tal Linzen. 2020 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
A systematic assessment of syntactic generalization in neural language models
Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger P Levy. 2020 · 2005
Earlier work this paper cites.
When BERT forgets how to POS: Amnesic probing of linguistic properties and MLM predictions
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2020 · 2006
Cited alongside, same era.
Learning phonology with substantive bias: An experimental and computational study of velar palatalization
Colin Wilson. 2006 · 2006
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2019 · 2019
Later among the works it cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Later among the works it cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstral, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li Li, and Razvan Pascanu. 2018 · 2018
Cited alongside, same era.
Neural models for reasoning over multiple mentions using coreference
Bhuwan Dhingra, Qiao Jin, Zhilin Yang, William Cohen, and Ruslan Salakhutdinov. 2018 · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Revisiting the poverty of the stimulus: hierarchical generalization without a hierarchical bias in recurrent neural networks
R Thomas McCoy, Robert Frank, and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Later among the works it cites.
Studying the inductive biases of rnns with synthetic variations of natural languages
Shauli Ravfogel, Yoav Goldberg, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Quantity doesn’t buy quality syntax with neural language models
Marten van Schijndel, Aaron Mueller, and Tal Linzen. 2019 · 2019
Later among the works it cites.
What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019 · 2019
Later among the works it cites.
Small and practical bert models for sequence labeling
Henry Tsai, Jason Riesa, Melvin Johnson, Naveen Arivazhagan, Xin Li, and Amelia Archer. 2019 · 2019
Later among the works it cites.
Does syntax need to grow on trees? sources of hierarchical inductive bias in sequence-to-sequence networks
R. Thomas McCoy, Robert Frank, and Tal Linzen. 2020 · 2020
Closest in time.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Closest in time.
Can neural networks acquire a structural bias from raw linguistic data?
Alex Warstadt and Samuel R Bowman. 2020 · 2020
Closest in time.
BLiMP: The benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.