Fetching the paper…
Reading the bibliography…
We identify empirical scaling laws for the cross-entropy loss in four domains: generative image modeling, video modeling, multimodal image$\leftrightarrow$text models, and mathematical problem solving.
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 1902
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli · 1904
Earlier work this paper cites.
One epoch is all you need, 2019, arXiv:1906.06669
Aran Komatsuzaki · 1906
Earlier work this paper cites.
Scaling autoregressive video models, 2019, 1906.02634
Dirk Weissenborn, Oscar Täckström, and Jakob Uszkoreit · 1906
Earlier work this paper cites.
Compositionality decomposed: how do neural networks generalise?, 2019, 1908.08351
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 1908
Earlier work this paper cites.
A constructive prediction of the generalization error across scales, 2019, 1909.12673
Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 1909
Earlier work this paper cites.
Imanol Schlag, Paul Smolensky, Roland Fernandez, Nebojsa Jojic, Jürgen Schmidhuber, and Jianfeng Gao · 1910
Earlier work this paper cites.
Scaling laws for neural language models, 2020, 2001.08361
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E. Gonzalez · 2002
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism, 2020, 2003.02218
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2003
Cited alongside, same era.
Recipes for building an open-domain chatbot, 2020, 2004.13637
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M. Smith, Y-Lan Boureau, and Jason Weston · 2004
Cited alongside, same era.
A neural scaling law from the dimension of the data manifold, 2020, 2004.10802
Utkarsh Sharma and Jared Kaplan · 2004
Cited alongside, same era.
Language models are few-shot learners, 2020, 2005.14165
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2017
Later among the works it cites.
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Generating wikipedia by summarizing long sequences
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Jukebox: A generative model for music, 2020, 2005.00341
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2005
Cited alongside, same era.
On the predictability of pruning across scales, 2020, 2006.10621
Jonathan S. Rosenfeld, Jonathan Frankle, Michael Carbin, and Nir Shavit · 2006
Cited alongside, same era.
The new data and new challenges in multimedia research
Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2015
Cited alongside, same era.
Pixel recurrent neural networks
Aäron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
A downsampled variant of imagenet as an alternative to the CIFAR datasets
Patryk Chrabaszcz, Ilya Loshchilov, and Frank Hutter · 2017
Cited alongside, same era.
Deep learning scaling is predictable, empirically, 2017, 1712.00409
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
Later among the works it cites.
An empirical model of large-batch training, 2018, arXiv:1812.06162
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Later among the works it cites.
Neural discrete representation learning, 2018, 1711.00937
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2018
Later among the works it cites.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Closest in time.