Fetching the paper…
Reading the bibliography…
The ever-growing diversity of pre-training text corpora has equipped language models with generalization capabilities across various downstream tasks.
Large text compression benchmark, 2011
Matt Mahoney · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Super-convergence: Very fast training of neural networks using large learning rates, 2018
Leslie N. Smith and Nicholay Topin · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Earlier work this paper cites.
Green ai, 2019
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Probabilistic active meta-learning
Jean Kaddour, Steindór Sæmundsson, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Glu variants improve transformer, 2020
Noam Shazeer · 2020
Earlier work this paper cites.
Jack Bandy and Nicholas Vincent · 2021
Earlier work this paper cites.
Benchmarking differential privacy and federated learning for bert models
Priyam Basu, Tiasa Singha Roy, Rakshit Naidu, Zumrut Muftuoglu, Sahib Singh, and Fatemehsadat Mireshghallah · 2021
Earlier work this paper cites.
Pitfalls in machine learning research: Reexamining the development cycle, 2021
Stella Biderman and Walter J. Scheirer · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
N Elhage, N Nanda, C Olsson, T Henighan, N Joseph, B Mann, A Askell, Y Bai, A Chen, T Conerly, et al · 2021
Cited alongside, same era.
How to train bert with an academic budget
Peter Izsak, Moshe Berchansky, and Omer Levy · 2021
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Cited alongside, same era.
Jurassic-1: Technical details and evaluation
Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham · 2021
Cited alongside, same era.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna · 2021
Cited alongside, same era.
Quality not quantity: On the interaction between dataset design and robustness of clip
Thao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh, and Ludwig Schmidt · 2022
Later among the works it cites.
In-context learning and induction heads, 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Language modelling with pixels
Phillip Rust, Jonas F Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S Morcos · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2022
Cited alongside, same era.
Cramming: Training a language model on a single gpu in one day
Jonas Geiping and Tom Goldstein · 2022
Cited alongside, same era.
Pile of law: Learning responsible data filtering from the law and a 256GB open-source legal dataset
Peter Henderson, Mark Simon Krass, Lucia Zheng, Neel Guha, Christopher D Manning, Dan Jurafsky, and Daniel E. Ho · 2022
Cited alongside, same era.
Scaling laws and interpretability of learning from repeated data
Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Nlp from scratch without large-scale pretraining: A simple and efficient framework
Xingcheng Yao, Yanan Zheng, Xiaocong Yang, and Zhilin Yang · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Morcos · 2023
Closest in time.
Eliciting latent predictions from transformers with the tuned lens, 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Closest in time.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster, 2023
Nolan Dey, Gurpreet Gosal, Zhiming, Chen, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, and Joel Hestness · 2023
Closest in time.
Hungry hungry hippos: Towards language modeling with state space models
Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Re · 2023
Closest in time.
Copy is all you need
Tian Lan, Deng Cai, Yan Wang, Heyan Huang, and Xian-Ling Mao · 2023
Closest in time.
Spawrious: A benchmark for fine control of spurious correlation biases
Aengus Lynch, Gbètondji JS Dovonon, Jean Kaddour, and Ricardo Silva · 2023
Closest in time.
MTEB Leaderboard
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2023
Closest in time.
nanoT5, 3 2023
Piotr Nawrot · 2023
Closest in time.
Call for papers – the babylm challenge: Sample-efficient pretraining on a developmentally plausible corpus, 2023
Alex Warstadt, Leshem Choshen, Aaron Mueller, Adina Williams, Ethan Wilcox, and Chengxu Zhuang · 2023
Closest in time.
Data selection for language models via importance resampling, 2023
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Closest in time.
Spikegpt: Generative pre-trained language model with spiking neural networks
Rui-Jie Zhu, Qihang Zhao, and Jason K Eshraghian · 2023
Closest in time.