Fetching the paper…
Reading the bibliography…
Predictions of word-by-word conditional probabilities from Transformer-based language models are often evaluated to model the incremental processing difficulty of human readers.
Transformer-based language model surprisal predicts human reading times best with about two billion training tokens
Byung-Doh Oh and William Schuler. 2023a · 1921
Earlier work this paper cites.
Foundations of the Theory of Probability
Andrey Nikolaevich Kolmogorov. 1933 · 1933
Earlier work this paper cites.
Eye movement control in reading: The role of word boundaries
Alexander Pollatsek and Keith Rayner. 1982 · 1982
Earlier work this paper cites.
Lexical guidance in human parsing: Locus and processing characteristics
Don C. Mitchell. 1987 · 1987
Earlier work this paper cites.
Ohio Supercomputer Center
Ohio Supercomputer Center. 1987 · 1987
Earlier work this paper cites.
Subcategorization and sentence processing
Paul Gorrell. 1991 · 1991
Earlier work this paper cites.
Unspaced text interferes with both word identification and eye movement control
Keith Rayner, Martin H. Fischer, and Alexander Pollatsek. 1998 · 1998
Earlier work this paper cites.
Space information is important for reading
Manuel Perea and Joana Acha. 2009 · 2000
Earlier work this paper cites.
A probabilistic Earley parser as a psycholinguistic model
John Hale. 2001 · 2001
Earlier work this paper cites.
The Dundee Corpus
Alan Kennedy, Robin Hill, and Joël Pynte. 2003 · 2003
Earlier work this paper cites.
Expectation-based syntactic comprehension
Roger Levy. 2008 · 2008
Cited alongside, same era.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
The Provo Corpus: A large eye-tracking corpus with predictability norms
Steven G. Luke and Kiel Christianson. 2018 · 2018
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023 · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Later among the works it cites.
Words, subwords, and morphemes: What really matters in the surprisal-reading time relationship?
Sathvik Nair and Philip Resnik. 2023 · 2023
Later among the works it cites.
Large GPT-like models are bad babies: A closer look at the relationship between linguistic competence and psycholinguistic measures
Julius Steuer, Marius Mosbach, and Dietrich Klakow. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
The Natural Stories corpus: A reading-time corpus of English texts containing rare syntactic constructions
Richard Futrell, Edward Gibson, Harry J. Tily, Idan Blank, Anastasia Vishnevetsky, Steven Piantadosi, and Evelina Fedorenko. 2021 · 2021
Cited alongside, same era.
Single-stage prediction models do not explain the magnitude of syntactic disambiguation difficulty
Marten van Schijndel and Tal Linzen. 2021 · 2021
Cited alongside, same era.
Syntactic surprisal from neural models predicts, but underestimates, human processing difficulty from syntactic ambiguities
Suhas Arehalli, Brian Dillon, and Tal Linzen. 2022 · 2022
Cited alongside, same era.
Why does surprisal from larger Transformer-based language models provide a poorer fit to human reading times?
Byung-Doh Oh and William Schuler. 2023b
Cited in the paper.
Testing the predictions of surprisal theory in 11 languages
Ethan Gotlieb Wilcox, Tiago Pimentel, Clara Meister, Ryan Cotterell, and Roger P. Levy. 2023 · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
Gemini: A family of highly capable multimodal models
Google Gemini Team. 2024 · 2024
Closest in time.
Large-scale benchmark yields no evidence that language model surprisal explains syntactic disambiguation difficulty
Kuan-Jung Huang, Suhas Arehalli, Mari Kugemoto, Christian Muxica, Grusha Prasad, Brian Dillon, and Tal Linzen. 2024 · 2024
Closest in time.
How to compute the probability of a word
Tiago Pimentel and Clara Meister. 2024 · 2024
Closest in time.
Large-scale evidence for logarithmic effects of word predictability on reading time
Cory Shain, Clara Meister, Tiago Pimentel, Ryan Cotterell, and Roger Levy. 2024 · 2024
Closest in time.