Fetching the paper…
Reading the bibliography…
Chomsky and others have very directly claimed that large language models (LLMs) are equally capable of learning languages that are possible and impossible for humans to learn.
Three models for the description of language
Noam Chomsky. 1956 · 1956
Earlier work this paper cites.
Syntactic Structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
On certain formal properties of grammars
Noam Chomsky. 1959 · 1959
Earlier work this paper cites.
Some universals of grammar with particular reference to the order of meaningful elements
Joseph Greenberg. 1963 · 1963
Earlier work this paper cites.
Aspects of the Theory of Syntax
Noam Chomsky. 1965 · 1965
Earlier work this paper cites.
Tree adjoining grammars: How much context-sensitivity is required to provide reasonable structural descriptions? , Studies in Natural Language Processing, page 206–250. Cambridge University Press
Aravind K. Joshi. 1985 · 1985
Earlier work this paper cites.
Evidence against the context-freeness of natural language
Stuart M. Shieber. 1985 · 1985
Earlier work this paper cites.
Language universals and linguistic typology: Syntax and morphology
Bernard Comrie. 1989 · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman. 1990 · 1990
Earlier work this paper cites.
A probabilistic Earley parser as a psycholinguistic model
John Hale. 2001 · 2001
Earlier work this paper cites.
On Nature and Language
Noam Chomsky. 2002 · 2002
Earlier work this paper cites.
The faculty of language: What is it, who has it, and how did it evolve?
Marc D. Hauser, Noam Chomsky, and W. Tecumseh Fitch. 2002 · 2002
Earlier work this paper cites.
Broca’s area and the language instinct
Mariacristina Musso, Andrea Moro, Volkmar Glauche, Michel Rijntjes, Jürgen Reichenbach, Christian Büchel, and Cornelius Weiller. 2003 · 2003
Earlier work this paper cites.
Constraints on multiple center-embedding of clauses
Fred Karlsson. 2007 · 2007
Earlier work this paper cites.
Expectation-based syntactic comprehension
Roger Levy. 2008 · 2008
Earlier work this paper cites.
The myth of language universals: Language diversity and its importance for cognitive science
Nicholas Evans and Stephen C Levinson. 2009 · 2009
Earlier work this paper cites.
RNNs can generate bounded hierarchical languages with optimal memory
John Hewitt, Michael Hahn, Surya Ganguli, Percy Liang, and Christopher D. Manning. 2020 · 2010
Earlier work this paper cites.
What does Pirahã grammar have to teach us about human language and the mind?
Daniel L. Everett. 2012 · 2012
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Context-free transductions with neural stacks
Yiding Hao, William Merrill, Dana Angluin, Robert Frank, Noah Amsel, Andrew Benz, and Simon Mendelsohn. 2018 · 2018
Earlier work this paper cites.
Depth-bounding is effective: Improvements and evaluation of unsupervised PCFG induction
Lifeng Jin, Finale Doshi-Velez, Timothy Miller, William Schuler, and Lane Schwartz. 2018 · 2018
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Cited alongside, same era.
What do RNN language models learn about filler–gap dependencies?
Ethan Wilcox, Roger Levy, Takashi Morita, and Richard Futrell. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Information-theoretic locality properties of natural language
Richard Futrell. 2019 · 2019
Cited alongside, same era.
Neural language models as psycholinguistic subjects: Representations of syntactic state
Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019 · 2019
Cited alongside, same era.
Sequential neural networks as automata
Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks?
Thang Pham, Trung Bui, Long Mai, and Anh Nguyen. 2021 · 2021
Later among the works it cites.
Attention is Turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic. 2021 · 2021
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Word order does matter and shuffled language models know it
Mostafa Abdou, Vinit Ravishankar, Artur Kulmizev, and Anders Søgaard. 2022 · 2022
Later among the works it cites.
Information theory as a bridge between language function and language form
Richard Futrell and Michael Hahn. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
William Merrill. 2019 · 2019
Cited alongside, same era.
Using priming to uncover the organization of syntactic representations in neural language models
Grusha Prasad, Marten van Schijndel, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
How can self-attention networks recognize Dyck-n languages?
Javid Ebrahimi, Dhruv Gelda, and Wei Zhang. 2020 · 2020
Cited alongside, same era.
Neural natural language inference models partially embed theories of lexical entailment and negation
Atticus Geiger, Kyle Richardson, and Christopher Potts. 2020 · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Cited alongside, same era.
Isabel Papadimitriou, Richard Futrell, and Kyle Mahowald. 2022 · 2022
Later among the works it cites.
Causal distillation for language models
Zhengxuan Wu, Atticus Geiger, Joshua Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, and Noah Goodman. 2022 · 2022
Later among the works it cites.
Conversations with Tyler: Noam Chomsky
Noam Chomsky. 2023 · 2023
Later among the works it cites.
Noam Chomsky: The false promise of ChatGPT
Noam Chomsky, Ian Roberts, and Jeffrey Watumull. 2023 · 2023
Later among the works it cites.
Neural networks and the Chomsky hierarchy
Gregoire Deletang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A Ortega. 2023 · 2023
Later among the works it cites.
What makes a language easy to deep-learn?
Lukas Galke, Yoav Ram, and Limor Raviv. 2023 · 2023
Later among the works it cites.
Qian Huang, Eric Zelikman, Sarah Li Chen, Yuhuai Wu, Gregory Valiant, and Percy Liang. 2023 · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. 2023 · 2023
Later among the works it cites.
The emergence of grammatical structure from inter-predictability
John Mansfield and Charles Kemp. 2023 · 2023
Later among the works it cites.
Large languages, impossible languages and human brains
Andrea Moro, Matteo Greco, and Stefano F. Cappa. 2023 · 2023
Later among the works it cites.
Pushdown layers: Encoding recursive structure in transformer language models
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D. Manning. 2023 · 2023
Later among the works it cites.
Injecting structural hints: Using language models to study inductive biases in language learning
Isabel Papadimitriou and Dan Jurafsky. 2023 · 2023
Later among the works it cites.
Alex Warstadt, Leshem Choshen, Aaron Mueller, Adina Williams, Ethan Wilcox, and Chengxu Zhuang. 2023 · 2023
Later among the works it cites.
Using computational models to test syntactic learnability
Ethan Gotlieb Wilcox, Richard Futrell, and Roger Levy. 2023 · 2023
Later among the works it cites.
Three reasons why AI doesn’t model human language
Johan J. Bolhuis, Stephen Crain, Sandiway Fong, and Andrea Moro. 2024 · 2024
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman. 2023 · 2024
Closest in time.
The Philosophy of Theoretical Linguistics: A Contemporary Outlook
Ryan M. Nefdt. 2024 · 2024
Closest in time.