Fetching the paper…
Reading the bibliography…
Current trends in pre-training Large Language Models (LLMs) primarily focus on the scaling of model and dataset size.
Task2vec: Task embedding for meta-learning
Alessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran, Subhransu Maji, Charless C. Fowlkes, Stefano Soatto, and Pietro Perona · 1902
Earlier work this paper cites.
Applied Bayesian Modeling and Causal Inference from Incomplete-Data Perspectives
Andrew Gelman and Xiao-Li Meng · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Exploring and predicting transferability across NLP tasks
Tu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, and Mohit Iyyer · 2005
Earlier work this paper cites.
Curse-of-dimensionality revisited: Collapse of the particle filter in very large scale systems , pp. 316–334
Thomas Bengtsson, Peter Bickel, and Bo Li · 2008
Earlier work this paper cites.
Obstacles to high-dimensional particle filtering
Chris Snyder, Thomas Bengtsson, Peter Bickel, and Jeff Anderson · 2008
Earlier work this paper cites.
Generative Language Modeling for Automated Theorem Proving
Stanislas Polu and Ilya Sutskever · 2009
Earlier work this paper cites.
Impossibility theorems for domain adaptation
S. B David, T Lu, T Luu, and D Pal · 2010
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The dynamic distance between learning tasks: From kolmogorov complexity to transfer learning via quantum physics and the information bottleneck of the weights of deep networks
Alessandro Achille, Glen Bigan Mbeng, and Stefano Soatto · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Insights on representational similarity in neural networks with canonical correlation
Ari S. Morcos, Maithra Raghu, and Samy Bengio · 2018
Earlier work this paper cites.
Assessing generative models via precision and recall, 2018
Mehdi S. M. Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly · 2018
Earlier work this paper cites.
Similarity of Neural Network Representations Revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Improved precision and recall metric for assessing generative models, 2019
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Cited alongside, same era.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Cited alongside, same era.
Revisiting precision and recall definition for generative model evaluation, 2019
Loïc Simon, Ryan Webster, and Julien Rabin · 2019
Cited alongside, same era.
Random network distillation as a diversity metric for both image and text generation, 2020
Liam Fowl, Micah Goldblum, Arjun Gupta, Amr Sharaf, and Tom Goldstein · 2020
Cited alongside, same era.
Unified scaling laws for routed language models
Aidan Clark, Diego De Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al · 2022
Later among the works it cites.
The vendi score: A diversity evaluation metric for machine learning, 2022
Dan Friedman and Adji Bousso Dieng · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Scaling laws for a multi-agent reinforcement learning model
Oren Neumann and Claudius Gros · 2022
Later among the works it cites.
Chinchilla’s wild implications
Nostalgebraist · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The pile: An 800gb dataset of diverse text for language modeling, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Reliable fidelity and diversity metrics for generative models
Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, and Jaejun Yoo · 2020
Cited alongside, same era.
The information complexity of learning tasks, their structure and their distance
Alessandro Achille, Giovanni Paolini, Glen Mbeng, and Stefano Soatto · 2021
Cited alongside, same era.
Data and parameter scaling laws for neural machine translation
Mitchell A Gordon, Kevin Duh, and Jared Kaplan · 2021
Cited alongside, same era.
Model performance scaling with multiple data sources
Tatsunori Hashimoto · 2021
Cited alongside, same era.
On the effect of pretraining corpora on in-context learning by a large-scale language model, 2022
Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woomyoung Park, Jung-Woo Ha, and Nako Sung · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference, 2022
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Later among the works it cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication, 2023
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S. Morcos · 2023
Closest in time.
Tinystories: How small can language models be and still speak coherent english?, 2023
Ronen Eldan and Yuanzhi Li · 2023
Closest in time.
Palm 2 technical report
Google · 2023
Closest in time.
Textbooks are all you need, 2023
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li · 2023
Closest in time.
S Longpre, G Yauney, E Reif, K Lee, A Roberts, B Zoph, D Zhou, J Wei, K Robinson, D Mimno, and D Ippolito · 2023
Closest in time.
Is pre-training truly better than meta-learning?
B. Miranda, P. Yu, S. Goyal, Y.-X. Wang, and S. Koyejo · 2023
Closest in time.
Measuring data, 2023
Margaret Mitchell, Alexandra Sasha Luccioni, Nathan Lambert, Marissa Gerchick, Angelina McMillan-Major, Ezinwanne Ozoani, Nazneen Rajani, Tristan Thrush, Yacine Jernite, and Douwe Kiela · 2023
Closest in time.
Gpt-4 technical report
OpenAI · 2023
Closest in time.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, and Nolan Dey · 2023
Closest in time.
Beyond neural scaling laws: beating power law scaling via data pruning, 2023
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos · 2023
Closest in time.
D4: Improving llm pretraining via document de-duplication and diversification, 2023
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari S. Morcos · 2023
Closest in time.
Effective pruning of web-scale datasets based on complexity of concept clusters, 2024
Amro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel, Kamalika Chaudhuri, and Ari S. Morcos · 2024
Closest in time.
Investigating data contamination for pre-training language models, 2024
Minhao Jiang, Ken Ziyu Liu, Ming Zhong, Rylan Schaeffer, Siru Ouyang, Jiawei Han, and Sanmi Koyejo · 2024
Closest in time.