Fetching the paper…
Reading the bibliography…
This paper introduces Filtered Corpus Training, a method that trains language models (LMs) on corpora with certain linguistic constructions filtered out from the training data, and uses it to measure the ability of LMs to perform linguistic generalization on the basis of indirect evidence.
Transformer-based language model surprisal predicts human reading times best with about two billion training tokens
Byung-Doh Oh and William Schuler. 2023a · 1921
Earlier work this paper cites.
Existential Sentences in English
Gary Milsark. 1974 · 1974
Earlier work this paper cites.
Lectures on Government and Binding
Noam Chomsky. 1993 · 1993
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A probabilistic Earley parser as a psycholinguistic model
John Hale. 2001 · 2001
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
On the predictive power of neural language models for human real-time comprehension behavior
Ethan Gotlieb Wilcox, Jon Gauthier, Jennifer Hu, Peng Qian, and Roger Philip Levy. 2020 · 2006
Earlier work this paper cites.
Expectation-based syntactic comprehension
Roger Levy. 2008 · 2008
Earlier work this paper cites.
Causality: Models, Reasoning and Inference , 2nd edition
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Modularity and intuitions in formal semantics: The case of polarity items
Emmanuel Chemla, Vincent Homer, and Daniel Rothschild. 2011 · 2011
Earlier work this paper cites.
Sequential vs. hierarchical syntactic models of human incremental sentence processing
Victoria Fossum and Roger Levy. 2012 · 2012
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Grammaticality, acceptability, and probability: A probabilistic view of linguistic knowledge
Jey Han Lau, Alexander Clark, and Shalom Lappin. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Universal Dependencies
Joakim Nivre, Daniel Zeman, Filip Ginter, and Francis Tyers. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Predictive power of word surprisal for reading times is a linear function of language model quality
Adam Goodkind and Klinton Bicknell. 2018 · 2018
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Do language models understand anything? on the ability of LSTMs to understand negative polarity items
Jaap Jumelet and Dieuwke Hupkes. 2018 · 2018
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
How can we ensure visibility and diversity in research contributions? How the Contributor Role Taxonomy (CRediT) is helping the shift from authorship to contributorship
Liz Allen, Alison O’Connell, and Veronique Kiermer. 2019 · 2019
Cited alongside, same era.
Neural language models as psycholinguistic subjects: Representations of syntactic state
Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019 · 2019
Cited alongside, same era.
BLiMP: The benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
The influence of polarity items on inferential judgments
Milica Denić, Vincent Homer, Daniel Rothschild, and Emmanuel Chemla. 2021 · 2021
Later among the works it cites.
Language models use monotonicity to assess NPI licensing
Jaap Jumelet, Milica Denic, Jakub Szymanik, Dieuwke Hupkes, and Shane Steinert-Threlkeld. 2021 · 2021
Later among the works it cites.
Lower perplexity is not always human-like
Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, and Kentaro Inui. 2021 · 2021
Later among the works it cites.
Refining targeted syntactic evaluation of language models
Benjamin Newman, Kai-Siang Ang, Julia Gong, and John Hewitt. 2021 · 2021
Later among the works it cites.
Frequency effects on syntactic rule learning in transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Verb argument structure alternations in word and sentence embeddings
Katharina Kann, Alex Warstadt, Adina Williams, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Using priming to uncover the organization of syntactic representations in neural language models
Grusha Prasad, Marten van Schijndel, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Investigating BERT’s knowledge of language: Five analysis methods with NPIs
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, and Samuel R. Bowman. 2019a · 2019
Cited alongside, same era.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Cited alongside, same era.
SyntaxGym: An online platform for targeted evaluation of language models
Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian, and Roger Levy. 2020 · 2020
Cited alongside, same era.
Probabilistic predictions of people perusing: Evaluating metrics of language model performance for psycholinguistic modeling
Yiding Hao, Simon Mendelsohn, Rachel Sterneck, Randi Martinez, and Robert Frank. 2020 · 2020
Cited alongside, same era.
Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick. 2021 · 2021
Later among the works it cites.
Syntactic surprisal from neural models predicts, but underestimates, human processing difficulty from syntactic ambiguities
Suhas Arehalli, Brian Dillon, and Tal Linzen. 2022 · 2022
Later among the works it cites.
Artificial Neural Networks as Models of Human Language Acquisition
Alex Warstadt. 2022 · 2022
Later among the works it cites.
What artificial neural networks can tell us about human language acquisition
Alex Warstadt and Samuel R. Bowman. 2022 · 2022
Later among the works it cites.
A taxonomy and review of generalization research in nlp
Dieuwke Hupkes, Mario Giulianelli, Verna Dankers, Mikel Artetxe, Yanai Elazar, Tiago Pimentel, Christos Christodoulopoulos, Karim Lasri, Naomi Saphra, Arabella Sinclair, Dennis Ulmer, Florian Schottmann, Khuyagbaatar Batsuren, Kaiser Sun, Koustuv Sinha, Leila Khalatbari, Maria Ryskina, Rita Frieske, Ryan Cotterell, and Zhijing Jin. 2023 · 2023
Later among the works it cites.
How to plant trees in language models: Data and architectural effects on the emergence of syntactic inductive biases
Aaron Mueller and Tal Linzen. 2023 · 2023
Later among the works it cites.
The validity of evaluation results: Assessing concurrence across compositionality benchmarks
Kaiser Sun, Adina Williams, and Dieuwke Hupkes. 2023 · 2023
Later among the works it cites.
Language models learn rare phenomena from less rare phenomena: The case of the missing aanns
Kanishka Misra and Kyle Mahowald. 2024 · 2024
Closest in time.
Frequency explains the inverse correlation of large language models’ size, training data amount, and surprisal’s fit to reading times
Byung-Doh Oh, Shisen Yue, and William Schuler. 2024 · 2024
Closest in time.
Interpretability of language models via task spaces
Lucas Weber, Jaap Jumelet, Elia Bruni, and Dieuwke Hupkes. 2024 · 2024
Closest in time.
Language modelling as a multi-task problem
Lucas Weber, Jaap Jumelet, Elia Bruni, and Dieuwke Hupkes. 2021 · 2060
Closest in time.