Fetching the paper…
Reading the bibliography…
Both humans and large language models are able to learn language without explicit structural supervision.
George Kingsley Zipf. 1936 · 1936
Earlier work this paper cites.
An informational theory of the statistical structure of language
Benoit Mandelbrot. 1953 · 1953
Earlier work this paper cites.
On certain formal properties of grammars
Noam Chomsky. 1959 · 1959
Earlier work this paper cites.
Tree adjoining grammars: How much context-sensitivity is required to provide reasonable structural descriptions?
Aravind K Joshi. 1985 · 1985
Earlier work this paper cites.
Evidence against the context-freeness of natural language
Stuart M. Shieber. 1985 · 1985
Earlier work this paper cites.
The convergence of mildly context-sensitive grammar formalisms
Aravind K Joshi, K Vijay Shanker, and David Weir. 1990 · 1990
Earlier work this paper cites.
Gapping as constituent coordination
Mark J. Steedman. 1990 · 1990
Earlier work this paper cites.
The minimalist program
Noam Chomsky. 1995 · 1995
Earlier work this paper cites.
Rethinking innateness: A connectionist perspective on development , volume 10
Jeffrey L Elman. 1996 · 1996
Earlier work this paper cites.
The faculty of language: What is it, who has it, and how did it evolve?
Marc D. Hauser, Noam Chomsky, and W. Tecumseh Fitch. 2002 · 2002
Earlier work this paper cites.
Children’s first language acquistion from a usage-based perspective
Elena Lieven and Michael Tomasello. 2008 · 2008
Earlier work this paper cites.
Computational perspectives on minimalism
Edward P Stabler. 2010 · 2010
Earlier work this paper cites.
Pre-training a language model without human language
Cheng-Han Chiang and Hung-yi Lee. 2020 · 2012
Earlier work this paper cites.
Statistical construction learning: Does a zipfian problem space ensure robust language learning
Nick C Ellis and Matthew Brook O’Donnell. 2012 · 2012
Earlier work this paper cites.
Zipf’s word frequency law in natural language: A critical review and future directions
Steven T Piantadosi. 2014 · 2014
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
The effect of zipfian frequency variations on category formation in adult artificial language learning
Kathryn D. Schuler, Patricia A. Reeder, Elissa L. Newport, and Richard N. Aslin. 2017 · 2017
Cited alongside, same era.
How efficiency shapes human language
Edward Gibson, Richard Futrell, Steven P Piantadosi, Isabelle Dautriche, Kyle Mahowald, Leon Bergen, and Roger Levy. 2019 · 2019
Cited alongside, same era.
Studying the inductive biases of RNNs with synthetic variations of natural languages
Shauli Ravfogel, Yoav Goldberg, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Universals of word order reflect optimization of grammars for efficient communication
Michael Hahn, Dan Jurafsky, and Richard Futrell. 2020 · 2020
Cited alongside, same era.
Learning music helps you read: Using transfer to study linguistic structure in language models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2020
Data distributional properties drive emergent in-context learning in transformers
Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill. 2022 · 2022
Later among the works it cites.
On the transferability of pre-trained language models: A study from artificial datasets
Cheng-Han Chiang and Hung-yi Lee. 2022 · 2022
Later among the works it cites.
The learnability consequences of zipfian distributions in language
Ori Lavi-Rotbain and Inbal Arnon. 2022 · 2022
Later among the works it cites.
Frozen pretrained transformers as universal computation engines
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch. 2022 · 2022
Later among the works it cites.
Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, and Sebastian Schuster. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning which features matter: RoBERTa acquires a preference for linguistic generalizations (eventually)
Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, and Samuel R. Bowman. 2020 · 2020
Cited alongside, same era.
Variation in mild context-sensitivity
Robert Frank and Tim Hunter. 2021 · 2021
Cited alongside, same era.
Initializing new word embeddings for pretrained language models
John Hewitt. 2021 · 2021
Cited alongside, same era.
What they do when in doubt: a study of inductive biases in seq2seq learners
Eugene Kharitonov and Rahma Chaabouni. 2021 · 2021
Cited alongside, same era.
Does pretraining for summarization require knowledge transfer?
Kundan Krishna, Jeffrey Bigham, and Zachary C. Lipton. 2021 · 2021
Cited alongside, same era.
Deep Learning and Linguistic Representation
Shalom Lappin. 2021 · 2021
Cited alongside, same era.
Structure here, bias there: Hierarchical generalization by jointly learning syntactic transformations
Karl Mulligan, Robert Frank, and Tal Linzen. 2021 · 2021
Cited alongside, same era.
Neural Network Approaches to the Study of Word Learning
Eva Portelance. 2022 · 2022
Later among the works it cites.
The effects of corpus choice and morphosyntax on multilingual space induction
Vinit Ravishankar and Joakim Nivre. 2022 · 2022
Later among the works it cites.
Pretraining with artificial language: Studying transferable knowledge in language models
Ryokan Ri and Yoshimasa Tsuruoka. 2022 · 2022
Later among the works it cites.
A Cognitive Bias for Zipfian Distributions? Uniform Distributions Become More Skewed via Cultural Transmission
Amir Shufaniya and Inbal Arnon. 2022 · 2022
Later among the works it cites.
What artificial neural networks can tell us about human language acquisition
A. Warstadt and Samuel R. Bowman. 2022 · 2022
Later among the works it cites.
Lexical semantic content, not syntactic structure, is the main contributor to ann-brain similarity of fMRI responses in the language network
Carina Kauf, Greta Tuckute, Roger Levy, Jacob Andreas, and Evelina Fedorenko. 2023 · 2023
Closest in time.
Characterizing intrinsic compositionality in transformers with tree projections
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D Manning. 2023 · 2023
Closest in time.
Alex Warstadt, Leshem Choshen, Aaron Mueller, Adina Williams, Ethan Wilcox, and Chengxu Zhuang. 2023 · 2023
Closest in time.
Using Computational Models to Test Syntactic Learnability
Ethan Gotlieb Wilcox, Richard Futrell, and Roger Levy. 2023 · 2023
Closest in time.
Oolong: Investigating what makes transfer learning hard with controlled studies
Zhengxuan Wu, Alex Tamkin, and Isabel Papadimitriou. 2023 · 2023
Closest in time.