Fetching the paper…
Reading the bibliography…
Transformers trained on natural language data have been shown to learn its hierarchical structure and generalize to sentences with unseen syntactic structures without explicitly encoding any structural bias.
A formal theory of inductive inference. part i
R.J. Solomonoff · 1964
Earlier work this paper cites.
Language and Mind
Noam Chomsky · 1968
Earlier work this paper cites.
Syntax directed translations and the pushdown assembler
Alfred V. Aho and Jeffrey D. Ullman · 1969
Earlier work this paper cites.
Modeling by shortest data description
J. Rissanen · 1978
Earlier work this paper cites.
Trainable grammars for speech recognition
James K. Baker · 1979
Earlier work this paper cites.
Language and Learning: The Debate Between Jean Piaget and Noam Chomsky
Noam Chomsky · 1980
Earlier work this paper cites.
The Philosophy and the Approach
David Marr · 1982
Earlier work this paper cites.
Inducing probabilistic grammars by bayesian model merging
Andreas Stolcke and Stephen Omohundro · 1994
Earlier work this paper cites.
Distributional information: A powerful cue for acquiring syntactic categories
Martin Redington, Nick Chater, and Steven Finch · 1998
Earlier work this paper cites.
The childes project: tools for analyzing talk
Brian Macwhinney · 2000
Earlier work this paper cites.
Learnability and the statistical structure of language: Poverty of stimulus arguments revisited
John D Lewis and Jeffrey L Elman · 2001
Earlier work this paper cites.
Structure dependence in language acquisition: Uncovering the statistical richness of the stimulus
Florencia Reali and Morten H Christiansen · 2004
Earlier work this paper cites.
Poverty of the stimulus? a rational approach
Amy Perfors, Terry Regier, and Joshua B Tenenbaum · 2006
Earlier work this paper cites.
Transformational networks
Robert Frank and Donald Mathis · 2007
Earlier work this paper cites.
Word learning as bayesian inference
Fei Xu and Joshua B Tenenbaum · 2007
Earlier work this paper cites.
A rational analysis of rule-based concept learning
Noah D Goodman, Joshua B Tenenbaum, Jacob Feldman, and Thomas L Griffiths · 2008
Earlier work this paper cites.
The learnability of abstract syntactic principles
Amy Perfors, Joshua B. Tenenbaum, and Terry Regier · 2010
Earlier work this paper cites.
Bayesian theory of mind: Modeling joint belief-desire attribution
Chris Baker, Rebecca Saxe, and Joshua Tenenbaum · 2011
Earlier work this paper cites.
Predicting pragmatic reasoning in language games
Michael C. Frank and Noah D. Goodman · 2012
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Brenden M. Lake, Ruslan Salakhutdinov, and Joshua B. Tenenbaum · 2015
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes, 2017
Guillaume Alain and Yoshua Bengio · 2017
Cited alongside, same era.
Improved neural machine translation with a syntax-aware encoder and decoder
Huadong Chen, Shujian Huang, David Chiang, and Jiajun Chen · 2017
Cited alongside, same era.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Tree-to-tree neural networks for program translation
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams · 2022
Later among the works it cites.
Grokking ‘grokking’, 2023
Beren Millidge · 2022
Later among the works it cites.
Coloring the blank slate: Pre-training imparts a hierarchical inductive bias to sequence-to-sequence models
Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, and Sebastian Schuster · 2022
Later among the works it cites.
Transformers can do bayesian inference
Samuel Müller, Noah Hollmann, Sebastian Pineda Arango, Josif Grabocka, and Frank Hutter · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harrison Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinyun Chen, Chang Liu, and Dawn Song · 2018
Cited alongside, same era.
Learning sparse neural networks through l_0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma · 2018
Cited alongside, same era.
R. Thomas McCoy, Roberta Frank, and Tal Linzen · 2018
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih · 2018
Cited alongside, same era.
Random deep neural networks are biased towards simple functions
Giacomo De Palma, Bobak Kiani, and Seth Lloyd · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2019
Cited alongside, same era.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon, 2022
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Later among the works it cites.
Physics of Language Models: Part 1, Learning Hierarchical Language Structures
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Later among the works it cites.
Simplicity bias in transformers and their ability to learn sparse Boolean functions
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom · 2023
Later among the works it cites.
A fine-grained comparison of pragmatic language understanding in humans and language models
Jennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko, and Edward Gibson · 2023
Later among the works it cites.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Later among the works it cites.
How to plant trees in language models: Data and architectural effects on the emergence of syntactic inductive biases
Aaron Mueller and Tal Linzen · 2023
Later among the works it cites.
In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax, November 2023
Aaron Mueller, Albert Webson, Jackson Petty, and Tal Linzen · 2023
Later among the works it cites.
Grokking of hierarchical structure in vanilla transformers
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher Manning · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability, January 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Arkil Patel, Satwik Bhattamishra, Siva Reddy, and Dzmitry Bahdanau · 2023
Later among the works it cites.
[AN #159]: Building agents that know how to experiment, by training on procedurally generated games
Rohin Shah · 2023
Later among the works it cites.
Clever hans or neural theory of mind? stress testing social reasoning in large language models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz · 2023
Later among the works it cites.
Explaining grokking through circuit efficiency, September 2023
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Later among the works it cites.
The heuristic core: Understanding subnetwork generalization in pretrained language models, 2024
Adithya Bhaskar, Dan Friedman, and Danqi Chen · 2024
Closest in time.
Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L Leavitt, and Naomi Saphra · 2024
Closest in time.