Fetching the paper…
Reading the bibliography…
Early in training, LMs can behave like n-gram models, but eventually they often learn tree-based syntactic rules and generalize hierarchically out of distribution (OOD).
Formal principles of language acquisition
Kenneth Wexler. 1980 · 1980
Earlier work this paper cites.
The child’s trigger experience: Degree-0 learnability
David Lightfoot. 1989 · 1989
Earlier work this paper cites.
Simple fast algorithms for the editing distance between trees and related problems
Kaizhong Zhang and Dennis Shasha. 1989 · 1989
Earlier work this paper cites.
Transformational networks
Robert Frank and Donald Mathis. 2007 · 2007
Earlier work this paper cites.
Poverty of the stimulus revisited
Robert C Berwick, Paul Pietroski, Beracah Yankama, and Noam Chomsky. 2011 · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
The childes project: Tools for analyzing talk, volume II: The database , 3 edition
Brian MacWhinney. 2014 · 2014
Earlier work this paper cites.
Aspects of the theory of syntax , 50 edition
Noam Chomsky. 2015 · 2015
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2016 · 2016
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
R Thomas McCoy, Robert Frank, and Tal Linzen. 2018 · 2018
Earlier work this paper cites.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
Understanding learning dynamics of language models with SVCCA
Naomi Saphra and Adam Lopez. 2018 · 2018
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Earlier work this paper cites.
Understanding learning dynamics of language models with
Naomi Saphra and Adam Lopez. 2019 · 2019
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020 · 2020
Earlier work this paper cites.
What shapes feature representations? exploring datasets, architectures, and training
Katherine L Hermann and Andrew K Lampinen. 2020 · 2020
Cited alongside, same era.
Learning music helps you read: Using transfer to study linguistic structure in language models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2020
Cited alongside, same era.
Benign interpolation of noise in deep learning
Marthinus Wilhelmus Theunissen, Marelie Davel, and Etienne Barnard. 2020 · 2020
Cited alongside, same era.
The curse of performance instability in analysis datasets: Consequences, source, and suggestions
Xiang Zhou, Yixin Nie, Hao Tan, and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
Transformers generalize linearly
Jackson Petty and Robert Frank. 2021 · 2021
Cited alongside, same era.
Latent state models of training dynamics
Michael Y Hu, Angelica Chen, Naomi Saphra, and Kyunghyun Cho. 2023 · 2023
Later among the works it cites.
ParaAMR: A large-scale syntactically diverse paraphrase dataset by AMR back-translation
Kuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar, Kai-Wei Chang, and Aram Galstyan. 2023 · 2023
Later among the works it cites.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks
William Merrill, Nikolaos Tsilivis, and Aman Shukla. 2023 · 2023
Later among the works it cites.
Aaron Mueller and Tal Linzen. 2023 · 2023
Later among the works it cites.
Grokking of hierarchical structure in vanilla transformers
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher Manning. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The MultiBERTs: BERT reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Raluca Turc, Jacob Eisenstein, Dipanjan Das, and Ellie Pavlick. 2021 · 2021
Cited alongside, same era.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang. 2022 · 2022
Cited alongside, same era.
The grammar-learning trajectories of neural language models
Leshem Choshen, Guy Hacohen, Daphna Weinshall, and Omri Abend. 2022 · 2022
Cited alongside, same era.
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, and 21 others. 2022 · 2022
Cited alongside, same era.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra. 2022 · 2022
Cited alongside, same era.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark. 2022 · 2022
Cited alongside, same era.
Coloring the blank slate: Pre-training imparts a hierarchical inductive bias to sequence-to-sequence models
Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, and Sebastian Schuster. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Injecting structural hints: Using language models to study inductive biases in language learning
Isabel Papadimitriou and Dan Jurafsky. 2023 · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. 2023 · 2023
Later among the works it cites.
Explaining grokking through circuit efficiency
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar. 2023 · 2023
Later among the works it cites.
Kabir Ahuja, Vidhisha Balachandran, Madhur Panwar, Tianxing He, Noah A Smith, Navin Goyal, and Yulia Tsvetkov. 2024 · 2024
Closest in time.
A dependency distance approach to the syntactic complexity variation in the connected speech of alzheimer’s disease
Nan Gao and Qingshun He. 2024 · 2024
Closest in time.
Yufei Huang, Shengding Hu, Xu Han, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Closest in time.
A percolation model of emergence: Analyzing transformers trained on a formal language
Ekdeep Singh Lubana, Kyogo Kawaguchi, Robert P Dick, and Hidenori Tanaka. 2024 · 2024
Closest in time.
Does syntax need to grow on trees? sources of hierarchical inductive bias in sequence-to-sequence networks
Tom McCoy, Robert Frank, and Tal Linzen. 2020b · 2024
Closest in time.
In-context learning generalizes, but not always robustly: The case of syntax
Aaron Mueller, Albert Webson, Jackson Petty, and Tal Linzen. 2024 · 2024
Closest in time.
Competition dynamics shape algorithmic phases of in-context learning
Core Francisco Park, Ekdeep Singh Lubana, Itamar Pres, and Hidenori Tanaka. 2024 · 2024
Closest in time.
Revisiting code similarity evaluation with abstract syntax tree edit distance
Yewei Song, Cedric Lothritz, Daniel Tang, Tegawendé F Bissyandé, and Jacques Klein. 2024 · 2024
Closest in time.
Critical data size of language models from a grokking perspective
Xuekai Zhu, Yao Fu, Bowen Zhou, and Zhouhan Lin. 2024 · 2024
Closest in time.