Fetching the paper…
Reading the bibliography…
We investigate what kind of structural knowledge learned in neural network encoders is transferable to processing natural language.
Human behavior and the principle of least effort
George Kingsley Zipf. 1949 · 1949
Earlier work this paper cites.
Syntactic Structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
Dependency Syntax: Theory and Practice
Igor Mel’čuk. 1988 · 1988
Earlier work this paper cites.
Word Association Norms, Mutual Information, and Lexicography
Kenneth Ward Church and Patrick Hanks. 1989 · 1989
Earlier work this paper cites.
Building a Large Annotated Corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Latent Dirichlet Allocation
David M. Blei, A. Ng, and Michael I. Jordan. 2003 · 2003
Earlier work this paper cites.
Dynamics of Text Generation with Realistic Zipf’s Distribution
Damián H. Zanette and Marcelo A. Montemurro. 2005 · 2005
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukás Burget, Jan Honza Cernocký, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Dependency-Based Word Embeddings
Omer Levy and Yoav Goldberg. 2014 · 2014
Earlier work this paper cites.
Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, and Tiago Luís. 2015 · 2015
Earlier work this paper cites.
A Latent Variable Model Approach to PMI-based Word Embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2016 · 2016
Earlier work this paper cites.
Deep Biaffine Attention for Neural Dependency Parsing
Timothy Dozat and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Earlier work this paper cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Visualisation and ’Diagnostic Classifiers’ Reveal how Recurrent and Recursive Neural Networks Process Hierarchical Structure (Extended Abstract)
Dieuwke Hupkes and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Cited alongside, same era.
LSTMs Exploit Linguistic Attributes of Data
Nelson F. Liu, Omer Levy, Roy Schwartz, Chenhao Tan, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
The Importance of Being Recurrent for Modeling Hierarchical Structure
Ke Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Cross-Lingual Ability of Multilingual BERT: An Empirical Study
Stephen Mayhew Karthikeyan K, Zihan Wang and Dan Roth. 2020 · 2020
Later among the works it cites.
Context Analysis for Pre-trained Masked Language Models
Yi-An Lai, Garima Lalwani, and Yi Zhang. 2020 · 2020
Later among the works it cites.
Multilingual Denoising Pre-training for Neural Machine Translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Later among the works it cites.
Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language Models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
How Multilingual is Multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
Studying the Inductive Biases of RNNs with Synthetic Variations of Natural Languages
Shauli Ravfogel, Yoav Goldberg, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT
Shijie Wu and Mark Dredze. 2019 · 2019
Cited alongside, same era.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Revisiting the Context Window for Cross-lingual Word Embeddings
Ryokan Ri and Yoshimasa Tsuruoka. 2020 · 2020
Later among the works it cites.
A Primer on Pretrained Multilingual Language Models
Sumanth Doddapaneni, Gowtham Ramesh, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh M. Khapra. 2021 · 2021
Later among the works it cites.
Pretrained Transformers as Universal Computation Engines
Kevin Lu, Aditya Grover, P. Abbeel, and Igor Mordatch. 2021 · 2021
Later among the works it cites.
Deep Subjecthood: Higher-Order Grammatical Features in Multilingual BERT
Isabel Papadimitriou, Ethan A. Chi, Richard Futrell, and Kyle Mahowald. 2021 · 2021
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021a · 2021
Later among the works it cites.
Examining the Inductive Bias of Neural Language Models with Artificial Languages
Jennifer C. White and Ryan Cotterell. 2021 · 2021
Later among the works it cites.
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
On the Transferability of Pre-trained Language Models: A Study from Artificial Datasets
Cheng-Han Chiang and Hung yi Lee. 2022 · 2022
Closest in time.
Can Wikipedia Help Offline Reinforcement Learning?
Machel Reid, Yutaro Yamada, and Shixiang Shane Gu. 2022 · 2022
Closest in time.