Fetching the paper…
Reading the bibliography…
One of the main challenges in optimal scaling of large language models (LLMs) is the prohibitive cost of hyperparameter tuning, particularly learning rate $\eta$ and batch size $B$.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Accurate, large minibatch sg d: training imagenet in 1 hour
P Goyal · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I Loshchilov · 2017
Earlier work this paper cites.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
SL Smith · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Earlier work this paper cites.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Daniel Park, Jascha Sohl-Dickstein, Quoc Le, and Samuel Smith · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2019
Earlier work this paper cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B Grosse · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
On the generalization benefit of noise in stochastic gradient descent
Samuel Smith, Erich Elsen, and Soham De · 2020
Earlier work this paper cites.
Feature learning in infinite-width neural networks
Greg Yang and Edward J Hu · 2020
Earlier work this paper cites.
Learning rates as a function of batch size: A random matrix theory approach to neural network training
Diego Granziol, Stefan Zohren, and Stephen Roberts · 2022
Cited alongside, same era.
Batch size-invariance for policy optimization
Jacob Hilton, Karl Cobbe, and John Schulman · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
On the sdes and scaling rules for adaptive gradient algorithms
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora · 2022
Cited alongside, same era.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward J Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2022
The practitioner’s guide to the maximal update parameterization, 2024
Nolan Dey, Quentin Anthony, and Joel Hestness · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Scaling exponents across parameterizations and optimizers
Katie Everett, Lechao Xiao, Mitchell Wortsman, Alexander A Alemi, Roman Novak, Peter J Liu, Izzeddin Gur, Jascha Sohl-Dickstein, Leslie Pack Kaelbling, Jaehoon Lee, et al · 2024
Closest in time.
Apple intelligence foundation language models
Tom Gunter, Zirui Wang, Chong Wang, Ruoming Pang, Andy Narayanan, Aonan Zhang, Bowen Zhang, Chen Chen, Chung-Cheng Chiu, David Qiu, et al · 2024
Closest in time.
Scaling laws and compute-optimal training beyond fixed training durations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Automatic gradient descent: Deep learning without hyperparameters
Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan, and Yisong Yue · 2023
Cited alongside, same era.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
Blake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin, and Cengiz Pehlevan · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Cited alongside, same era.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Cited alongside, same era.
Beyond implicit bias: The insignificance of sgd noise in online learning
Nikhil Vyas, Depen Morwani, Rosie Zhao, Gal Kaplun, Sham Kakade, and Boaz Barak · 2023
Cited alongside, same era.
Alexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro Von Werra, and Martin Jaggi · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al · 2024
Closest in time.
Simple and scalable strategies to continually pre-train large language models
Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish · 2024
Closest in time.
Analyzing & eliminating learning rate warmup in GPT pre-training
Atli Kosson, Bettina Messmer, and Martin Jaggi · 2024
Closest in time.
Scalable optimization in the modular norm
Tim Large, Yang Liu, Minyoung Huh, Hyojin Bahng, Phillip Isola, and Jeremy Bernstein · 2024
Closest in time.
Surge phenomenon in optimal learning rate and batch size scaling
Shuaipeng Li, Penghao Zhao, Hailin Zhang, Xingwu Sun, Hao Wu, Dian Jiao, Weiyan Wang, Chengjun Liu, Zheng Fang, Jinbao Xue, et al · 2024
Closest in time.
A large-scale exploration of μ \mu -transfer
Lucas Lingle · 2024
Closest in time.
Llm foundry, 2024
MosaicML · 2024
Closest in time.
Power scheduler: A batch size and token number agnostic learning rate scheduler
Yikang Shen, Matthew Stallone, Mayank Mishra, Gaoyuan Zhang, Shawn Tan, Aditya Prasad, Adriana Meza Soria, David D Cox, and Rameswar Panda · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Investigating continual pretraining in large language models: Insights and implications
Çağatay Yıldız, Nishaanth Kanna Ravichandran, Prishruit Punia, Matthias Bethge, and Beyza Ermis · 2024
Closest in time.