Fetching the paper…
Reading the bibliography…
How much data is required to learn the structure of a language via next-token prediction? We study this question for synthetic datasets generated via a Probabilistic Context-Free Grammar (PCFG) -- a tree-like generative model that captures many of the hierarchical structures found in natural languages.
Aspects of the Theory of Syntax
Noam Chomsky · 1965
Earlier work this paper cites.
A Study of Grammatical Inference
J.J. Horning · 1969
Earlier work this paper cites.
Evidence against the context-freeness of natural language
S.M. Shieber · 1985
Earlier work this paper cites.
Statistical learning by 8-month-old infants
Jenny R Saffran, Richard N Aslin, and Elissa L Newport · 1996
Earlier work this paper cites.
Handbook of Formal Languages
Grzegorz Rozenberg and Arto Salomaa · 1997
Earlier work this paper cites.
The use of predictive dependencies in language learning
Jenny R Saffran · 2001
Earlier work this paper cites.
Frequency effects in language processing: A review with implications for theories of implicit and explicit language acquisition
Nick C Ellis · 2002
Earlier work this paper cites.
Pac-learning unambiguous nts languages
Alexander Clark · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Poverty of the stimulus revisited
Robert C Berwick, Paul Pietroski, Beracah Yankama, and Noam Chomsky · 2011
Earlier work this paper cites.
Formal language theory: refining the chomsky hierarchy
Gerhard Jäger and James Rogers · 2012
Earlier work this paper cites.
The unreasonable effectiveness of recurrent neural networks, 2015
2015
Earlier work this paper cites.
Deep learning and hierarchal generative models
E. Mossel · 2016
Earlier work this paper cites.
Critical behavior in physics and probabilistic formal languages
Henry W. Lin and Max Tegmark · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Infant statistical learning
Jenny R Saffran and Natasha Z Kirkham · 2018
Earlier work this paper cites.
Improving language understanding with unsupervised learning
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Dissecting contextual word embeddings: Architecture and representation
M. E. Peters, M. Neumann, L. Zettlemoyer, and W. Yih · 2018
Earlier work this paper cites.
A provably correct algorithm for deep learning that actually works
Eran Malach and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
I. Tenney, D. Das, and E. Pavlick · 2019
Cited alongside, same era.
Random language model
E. DeGiuli · 2019
Cited alongside, same era.
Emergence of order in random languages
Eric De Giuli · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Later among the works it cites.
Physics of language models: Part 1, context-free grammar
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Later among the works it cites.
Do transformers parse while predicting the masked word?
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, and Sanjeev Arora · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Parallels in the sequential organization of birdsong and human speech
Tim Sainburg, Brad Theilman, Marvin Thielk, and Timothy Q Gentner · 2019
Cited alongside, same era.
Emergent linguistic structure in artificial neural networks trained by self-supervision
C. D Manning, K. Clark, J. Hewitt, U. Khandelwal, and O. Levy · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
The implications of local correlation on learning some deep functions
E. Malach and S. Shalev-Shwartz · 2020
Cited alongside, same era.
Consistent unsupervised estimators for anchored PCFGs
Alexander Clark and Nathanaël Fijalkow · 2020
Cited alongside, same era.
Does syntax need to grow on trees? sources of hierarchical inductive bias in sequence-to-sequence networks
R. Thomas McCoy, Robert Frank, and Tal Linzen · 2020
Cited alongside, same era.
Sanjeev Arora and Anirudh Goyal · 2023
Later among the works it cites.
Michael R Douglas · 2023
Later among the works it cites.
Autocorrelations decay in texts and applicability limits of language models
Nikolay Mikhaylovskiy and Ilya Churilov · 2023
Later among the works it cites.
What can be learnt with wide convolutional neural networks?
Francesco Cagnetta, Alessandro Favero, and Matthieu Wyart · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2024
Closest in time.
How deep neural networks learn compositional data: The random hierarchy model
Francesco Cagnetta, Leonardo Petrini, Umberto M. Tomasini, Alessandro Favero, and Matthieu Wyart · 2024
Closest in time.
How deep networks learn sparse and hierarchical data: the sparse random hierarchy model
Umberto Tomasini and Matthieu Wyart · 2024
Closest in time.
A phase transition in diffusion models reveals the hierarchical nature of data
Antonio Sclocchi, Alessandro Favero, and Matthieu Wyart · 2024
Closest in time.
Song Mei · 2024
Closest in time.
Kabir Ahuja, Vidhisha Balachandran, Madhur Panwar, Tianxing He, Noah A Smith, Navin Goyal, and Yulia Tsvetkov · 2024
Closest in time.
A distributional simplicity bias in the learning dynamics of transformers
Riccardo Rende, Federica Gerace, Alessandro Laio, and Sebastian Goldt · 2024
Closest in time.
Critical phase transition in a large language model
Kai Nakaishi, Yoshihiko Nishikawa, and Koji Hukushima · 2024
Closest in time.
A dynamical model of neural scaling laws
Blake Bordelon, Alexander Atanasov, and Cengiz Pehlevan · 2024
Closest in time.
Yatin Dandi, Emanuele Troiani, Luca Arnaboldi, Luca Pesce, Lenka Zdeborová, and Florent Krzakala · 2024
Closest in time.