Fetching the paper…
Reading the bibliography…
Are $n$-gram language models still relevant in this era of neural large language models (LLMs)? Our answer is yes, and we showcase their values in both text analysis and improving neural LLMs.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
Slava M. Katz · 1987
Earlier work this paper cites.
Speech and language processing - an introduction to natural language processing, computational linguistics, and speech recognition
Dan Jurafsky and James H. Martin · 2000
Earlier work this paper cites.
All our n-gram are belong to you
Alex Franz and Thorsten Brants · 2006
Earlier work this paper cites.
Linear work suffix array construction
Juha Kärkkäinen, Peter Sanders, and Stefan Burkhardt · 2006
Earlier work this paper cites.
Large language models in machine translation
T. Brants, Ashok Popat, Peng Xu, Franz Josef Och, and Jeffrey Dean · 2007
Earlier work this paper cites.
The infinite markov model
Daichi Mochihashi and Eiichiro Sumita · 2007
Earlier work this paper cites.
A stochastic memoizer for sequence data
Frank D. Wood, C. Archambeau, Jan Gasthaus, Lancelot F. James, and Yee Whye Teh · 2009
Earlier work this paper cites.
Using suffix arrays as language models: Scaling the n-gram
Herman Stehouwer and Menno van Zaanen · 2010
Earlier work this paper cites.
Quantitative analysis of culture using millions of digitized books
Erez Lieberman Aiden and Jean-Baptiste Michel · 2011
Earlier work this paper cites.
Suffix trees as language models
Casey Redd Kennington, Martin Kay, and Annemarie Friedrich · 2012
Earlier work this paper cites.
Context dependent recurrent neural network language model
Tomas Mikolov and Geoffrey Zweig · 2012
Earlier work this paper cites.
Mining source code repositories at massive scale using language modeling
Miltiadis Allamanis and Charles Sutton · 2013
Earlier work this paper cites.
Compact, efficient and unlimited capacity: Language modeling with compressed suffix trees
Ehsan Shareghi, Matthias Petri, Gholamreza Haffari, and Trevor Cohn · 2015
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Kilograms: Very large n-grams for malware classification
Edward Raff, William Fleming, Richard Zak, H. Anderson, Bill Finlayson, Charles K. Nicholas, and Mark McLean · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
Acl 2023 tutorial: Retrieval-based language models and applications
Akari Asai, Sewon Min, Zexuan Zhong, and Danqi Chen · 2023
Later among the works it cites.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, L. Sifre, and John M. Jumper · 2023
Later among the works it cites.
Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hanna Hajishirzi, Noah A. Smith, and Jesse Dodge · 2023
Later among the works it cites.
Openllama: An open reproduction of llama, May 2023
Xinyang Geng and Hao Liu · 2023
Later among the works it cites.
The big friendly filter
Dirk Groeneveld · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Cited alongside, same era.
Retrieval augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang · 2020
Cited alongside, same era.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Ana Marasovic, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2022
Cited alongside, same era.
Few-shot learning with retrieval augmented language models
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave · 2022
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2022
Cited alongside, same era.
Rest: Retrieval-based speculative decoding
Zhenyu He, Zexuan Zhong, Tianle Cai, Jason D Lee, and Di He · 2023
Later among the works it cites.
Copy is all you need
Tian Lan, Deng Cai, Yan Wang, Heyan Huang, and Xian-Ling Mao · 2023
Later among the works it cites.
Data portraits: Recording foundation model training data
Marc Marone and Benjamin Van Durme · 2023
Later among the works it cites.
The roots search tool: Data transparency for llms
Aleksandra Piktus, Christopher Akiki, Paulo Villegas, Hugo Laurenccon, Gérard Dupont, Alexandra Sasha Luccioni, Yacine Jernite, and Anna Rogers · 2023
Later among the works it cites.
REPLUG: Retrieval-augmented black-box language models
Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih · 2023
Later among the works it cites.
Dolma: An Open Corpus of 3 Trillion Tokens for Language Model Pretraining Research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Khyathi Chandu, Jennifer Dumas, Li Lucy, Xinxi Lyu, Ian Magnusson, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Pete Walsh, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo · 2023
Later among the works it cites.
RedPajama: An open source recipe to reproduce LLaMA training dataset, 2023
Together · 2023
Later among the works it cites.
Koala: An index for quantifying overlaps with pre-training corpora
Thuy-Trang Vu, Xuanli He, Gholamreza Haffari, and Ehsan Shareghi · 2023
Later among the works it cites.
Reliable, adaptable, and attributable language models with retrieval
Akari Asai, Zexuan Zhong, Danqi Chen, Pang Wei Koh, Luke Zettlemoyer, Hanna Hajishirzi, and Wen tau Yih · 2024
Closest in time.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Closest in time.