Fetching the paper…
Reading the bibliography…
How do language models learn to make predictions during pre-training? To study this, we extract learning curves from five autoregressive English language model pre-training runs, for 1M unseen tokens in context.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
Slava Katz. 1987 · 1987
Earlier work this paper cites.
Latent Dirichlet allocation
David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003 · 2003
Earlier work this paper cites.
Expectation-based syntactic comprehension
Roger Levy. 2008 · 2008
Earlier work this paper cites.
The associative structure of language: Contextual diversity in early word learning
Thomas Hills, Josita Maouene, Brian Riordan, and Linda Smith. 2010 · 2010
Earlier work this paper cites.
The influence of contextual diversity on word learning
Brendan Johns, Melody Dye, and Michael Jones. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Generalized Additive Models: An Introduction with R
Simon Wood. 2017 · 2017
Earlier work this paper cites.
Predictive power of word surprisal for reading times is a linear function of language model quality
Adam Goodkind and Klinton Bicknell. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
pyGAM: Generalized additive models in Python
Daniel Servén and Charlie Brummitt. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Neural language models as psycholinguistic subjects: Representations of syntactic state
Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Understanding learning dynamics of language models with SVCCA
Naomi Saphra and Adam Lopez. 2019 · 2019
Earlier work this paper cites.
Pretrained language model embryology: The birth of ALBERT
Cheng-Han Chiang, Sung-Feng Huang, and Hung-yi Lee. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
spaCy: Industrial-strength natural language processing in python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Cited alongside, same era.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Benjamin Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom B. Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022 · 2022
Later among the works it cites.
ChatGPT: Optimizing language models for dialogue
OpenAI. 2022 · 2022
Later among the works it cites.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022 · 2022
Later among the works it cites.
Contextual diversity favors the learning of new words in children regardless of their comprehension skills
Eva Rosa, Rafael Salom, and Manuel Perea. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Harnessing the power of LLMs in practice: A survey on ChatGPT and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Shaochen Zhong, Bing Yin, and Xia Hu. 2024 · 2020
Cited alongside, same era.
Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus
Julien Abadji, Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2021 · 2021
Cited alongside, same era.
How is BERT surprised? layerwise detection of linguistic anomalies
Bai Li, Zining Zhu, Guillaume Thomas, Yang Xu, and Frank Rudzicz. 2021 · 2021
Cited alongside, same era.
Probing across time: What does RoBERTa know and when?
Zeyu Liu, Yizhong Wang, Jungo Kasai, Hannaneh Hajishirzi, and Noah A. Smith. 2021 · 2021
Cited alongside, same era.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
Frequency effects on syntactic rule learning in transformers
Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick. 2021 · 2021
Cited alongside, same era.
Analyzing the mono- and cross-lingual pretraining dynamics of multilingual language models
Terra Blevins, Hila Gonen, and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, et al. 2022 · 2022
Later among the works it cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
Kushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022 · 2022
Later among the works it cites.
Introducing Claude
Anthropic. 2023 · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023 · 2023
Closest in time.
Language acquisition: Do children and language models follow similar learning stages?
Linnea Evanson, Yair Lakretz, and Jean Rémi King. 2023 · 2023
Closest in time.
PaLM 2 technical report
Google. 2023 · 2023
Closest in time.
Dissociating language and thought in large language models: A cognitive perspective
Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2023 · 2023
Closest in time.
Why does surprisal from larger transformer-based language models provide a poorer fit to human reading times?
Byung-Doh Oh and William Schuler. 2023 · 2023
Closest in time.
Alex Warstadt and Samuel R. Bowman. 2023 · 2023
Closest in time.
Training trajectories of language models across scales
Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, and Veselin Stoyanov. 2023 · 2023
Closest in time.
Strong Prediction: Language Model Surprisal Explains Multiple N400 Effects
James A. Michaelov, Megan D. Bardolph, Cyma K. Van Petten, Benjamin K. Bergen, and Seana Coulson. 2024 · 2024
Closest in time.