Fetching the paper…
Reading the bibliography…
Language models (LMs) estimate a probability distribution over strings in a natural language; these distributions are crucial for computing perplexity and surprisal in linguistics research.
Transformer-based language model surprisal predicts human reading times best with about two billion training tokens
Byung-Doh Oh and William Schuler. 2023a · 1921
Earlier work this paper cites.
The Psychobiology of Language
George K. Zipf. 1935 · 1935
Earlier work this paper cites.
Zipf’s law and Miller’s random-monkey model
Davis Howes. 1968 · 1968
Earlier work this paper cites.
Paradigms and processes in reading comprehension
Marcel Adam Just, Patricia A. Carpenter, and Jacqueline D. Woolley. 1982 · 1982
Earlier work this paper cites.
Studies on the phonological word
Tracy Alan Hall and Ursula Kleinhenz. 1999 · 1999
Earlier work this paper cites.
A probabilistic Earley parser as a psycholinguistic model
John Hale. 2001 · 2001
Earlier work this paper cites.
Word: a typological framework
Robert M. W. Dixon and Alexandra Y. Aikhenvald. 2003 · 2003
Earlier work this paper cites.
The Dundee corpus
Alan Kennedy, Robin Hill, and Joel Pynte. 2003 · 2003
Earlier work this paper cites.
Elements of Information Theory , second edition
Thomas M. Cover and Joy A. Thomas. 2006 · 2006
Earlier work this paper cites.
Speakers optimize information density through syntactic reduction
Roger Levy and T. Florian Jaeger. 2007 · 2007
Earlier work this paper cites.
Prosodic Phonology: With a New Foreword
Marina Nespor and Irene Vogel. 2007 · 2007
Earlier work this paper cites.
Expectation-based syntactic comprehension
Roger Levy. 2008 · 2008
Earlier work this paper cites.
Optimal processing times in reading: a formal model and empirical investigation
Nathaniel J. Smith and Roger Levy. 2008 · 2008
Earlier work this paper cites.
Word lengths are optimized for efficient communication
Steven T. Piantadosi, Harry Tily, and Edward Gibson. 2011 · 2011
Earlier work this paper cites.
The effect of word predictability on reading time is logarithmic
Nathaniel J. Smith and Roger Levy. 2013 · 2013
Earlier work this paper cites.
Zipf’s law of abbreviation as a language universal
Christian Bentz and Ramon Ferrer-i-Cancho. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
The natural stories corpus
Richard Futrell, Edward Gibson, Harry J. Tily, Idan Blank, Anastasia Vishnevetsky, Steven Piantadosi, and Evelina Fedorenko. 2018 · 2018
Cited alongside, same era.
Predictive power of word surprisal for reading times is a linear function of language model quality
Adam Goodkind and Klinton Bicknell. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
The Provo corpus: A large eye-tracking corpus with predictability norms
Between words and characters: A brief history of open-vocabulary modeling and tokenization in NLP
Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan. 2021 · 2021
Later among the works it cites.
Frequency, informativity and word length: Insights from typologically diverse corpora
Natalia Levshina. 2022 · 2022
Later among the works it cites.
Comparison of structural parsers and neural language models as surprisal estimators
Byung-Doh Oh, Christian Clark, and William Schuler. 2022 · 2022
Later among the works it cites.
mGPT: Few-shot learners go multilingual
Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Vladislav Mikhailov, Anastasia Kozlova, and Tatiana Shavrina. 2022 · 2022
Later among the works it cites.
Expanding horizons of cross-linguistic research on reading: The multilingual eye-movement corpus (MECO)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Steven G. Luke and Kiel Christianson. 2018 · 2018
Cited alongside, same era.
Investigating the effectiveness of BPE: The power of shorter sequences
Matthias Gallé. 2019 · 2019
Cited alongside, same era.
How efficiency shapes human language
Edward Gibson, Richard Futrell, Steven T. Piantadosi, Isabelle Dautriche, Kyle Mahowald, Leon Bergen, and Roger Levy. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
A large-scale study of the effects of word frequency and predictability in naturalistic reading
Cory Shain. 2019 · 2019
Cited alongside, same era.
Wiki-40B: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou. 2020 · 2020
Cited alongside, same era.
On the predictive power of neural language models for human real-time comprehension behavior
Ethan Wilcox, Jon Gauthier, Jennifer Hu, Peng Qian, and Roger Levy. 2020 · 2020
Cited alongside, same era.
Noam Siegelman, Sascha Schroeder, Cengiz Acartürk, Hee-Don Ahn, Svetlana Alexeeva, Simona Amenta, Raymond Bertram, Rolando Bonandrini, Marc Brysbaert, Daria Chernova, Sara Maria Da Fonseca, Nicolas Dirix, Wouter Duyck, Argyro Fella, Ram Frost, Carolina A. Gattei, Areti Kalaitzi, Nayoung Kwon, Kaidi Lõo, Marco Marelli, Timothy C. Papadopoulos, Athanassios Protopapas, Satu Savo, Diego E. Shalom, Natalia Slioussar, Roni Stein, Longjiao Sui, Analí Taboh, Veronica Tønnesen, Kerem Alp Usal, and Victor Kuperman. 2022 · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023 · 2023
Later among the works it cites.
A measure-theoretic characterization of tight language models
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Defining the word
Martin Haspelmath. 2023 · 2023
Later among the works it cites.
Revisiting the optimality of word lengths
Tiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald, and Ryan Cotterell. 2023a · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
Language model quality correlates with psychometric predictive power in multiple languages
Ethan Wilcox, Clara Meister, Ryan Cotterell, and Tiago Pimentel. 2023a · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
On the proper treatment of tokenization in psycholinguistics
Mario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell, Tim Vieira, and Ryan Cotterell. 2024 · 2024
Closest in time.
Byung-Doh Oh and William Schuler. 2024 · 2024
Closest in time.
Large-scale evidence for logarithmic effects of word predictability on reading time
Cory Shain, Clara Meister, Tiago Pimentel, Ryan Cotterell, and Roger Levy. 2024 · 2024
Closest in time.