Fetching the paper…
Reading the bibliography…
For humans, language production and comprehension is sensitive to the hierarchical structure of sentences.
Aspects of the Theory of Syntax
Noam Chomsky. 1965 · 1965
Earlier work this paper cites.
Structure dependence in grammar formation
Stephen Crain and Mineharu Nakayama. 1987 · 1987
Earlier work this paper cites.
A procedure for quantitatively comparing the syntactic coverage of English grammars
E. Black, S. Abney, D. Flickenger, C. Gdaniec, R. Grishman, P. Harrison, D. Hindle, R. Ingria, F. Jelinek, J. Klavans, M. Liberman, M. Marcus, S. Roukos, B. Santorini, and T. Strzalkowski. 1991 · 1991
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. 2018 · 2018
Earlier work this paper cites.
The importance of being recurrent for modeling hierarchical structure
Ke M Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Cited alongside, same era.
Does syntax need to grow on trees? Sources of hierarchical inductive bias in sequence-to-sequence networks
R. Thomas McCoy, Robert Frank, and Tal Linzen. 2020 · 2020
Cited alongside, same era.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Juergen Schmidhuber. 2021 · 2021
Cited alongside, same era.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz, and Noah A. Smith. 2021 · 2021
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022 · 2022
Later among the works it cites.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark. 2022 · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A. Smith. 2022 · 2022
Later among the works it cites.
Coloring the blank slate: Pre-training imparts a hierarchical inductive bias to sequence-to-sequence models
Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, and Sebastian Schuster. 2022 · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jackson Petty and Robert Frank. 2021 · 2021
Cited alongside, same era.
Characterizing intrinsic compositionality in transformers with tree projections
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D Manning. 2023 · 2023
Closest in time.