Fetching the paper…
Reading the bibliography…
We study empirical scaling laws for language model performance on the cross-entropy loss.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 1901
Earlier work this paper cites.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 1901
Earlier work this paper cites.
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 1902
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 1905
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems, 2019, 1905.00537
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 1905
Earlier work this paper cites.
One epoch is all you need, 2019, arXiv:1906.06669
Aran Komatsuzaki · 1906
Earlier work this paper cites.
Autogrow: Automatic layer growing in deep convolutional networks, 2019, 1906.02909
Wei Wen, Feng Yan, and Hai Li · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding, 2019, arXiv:1906.08237
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 1906
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George E. Dahl, Christopher J. Shallue, and Roger B. Grosse · 1907
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations, 2019, 1909.11942
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 1909
Earlier work this paper cites.
A constructive prediction of the generalization error across scales, 2019, 1909.12673
Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 1909
Earlier work this paper cites.
A constructive prediction of the generalization error across scales, 2019, arXiv:1909.12673
Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 1909
Earlier work this paper cites.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
Entropy and long-range correlations in literary english
Werner Ebeling and Thorsten Pöschel · 1994
Earlier work this paper cites.
Scaling to very very large corpora for natural language disambiguation
Michele Banko and Eric Brill · 2001
Earlier work this paper cites.
A bit of progress in language modeling
Joshua Goodman · 2001
Cited alongside, same era.
All of nonparametric statistics
Larry Wasserman · 2006
Cited alongside, same era.
On the origin of long-range correlations in texts
Eduardo G Altmann, Giampaolo Cristadoro, and Mirko Degli Esposti · 2012
Cited alongside, same era.
Analysis of a random forests model
GÊrard Biau · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Cited alongside, same era.
Adam: A method for stochastic optimization, 2014, 1412.6980
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Reconciling modern machine learning and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Later among the works it cites.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018
Later among the works it cites.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer · 2018
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
Criticality in formal languages and statistical physics
Henry W Lin and Max Tegmark · 2016
Cited alongside, same era.
Residual networks behave like ensembles of relatively shallow networks, 2016, arXiv:1605.06431
Andreas Veit, Michael Wilber, and Serge Belongie · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani and Andrew M. Saxe · 2017
Cited alongside, same era.
Yanping Huang, Yonglong Cheng, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, and Zhifeng Chen · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Generating wikipedia by summarizing long sequences
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer · 2018
Later among the works it cites.
An empirical model of large-batch training, 2018, arXiv:1812.06162
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Later among the works it cites.
The full spectrum of deep net hessians at scale: Dynamics with sample size
Vardan Papyan · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
Mesh-tensorflow: Deep learning for supercomputers, 2018, 1811.02084
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, and Blake Hechtman · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training, 2018, arXiv:1811.03600
Christopher J. Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Later among the works it cites.
Introduction to the theory of complex systems
Stefan Thurner, Rudolf Hanel, and Peter Klimek · 2018
Later among the works it cites.
Beyond human-level accuracy: Computational challenges in deep learning
Joel Hestness, Newsha Ardalani, and Gregory Diamos · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.