Fetching the paper…
Reading the bibliography…
Training large models is both resource-intensive and time-consuming, making it crucial to understand the quantitative relationship between model performance and hyperparameters.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
Practical Recommendations for Gradient-Based Training of Deep Architectures , pp. 437–478
Yoshua Bengio · 2012
Earlier work this paper cites.
Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves
Tobias Domhan, Jost Tobias Springenberg, and Frank Hutter · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, et al · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, et al · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, et al · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Learning curve prediction with bayesian neural networks
Aaron Klein, Stefan Falkner, Jost Tobias Springenberg, and Frank Hutter · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, et al · 2017
Earlier work this paper cites.
On the difficulty of dnn hyperparameter optimization using learning curve prediction
Daeyoung Choi, Hyunghun Cho, and Wonjong Rhee · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, et al · 2019
Earlier work this paper cites.
Learning an adaptive learning rate schedule
Zhen Xu, Andrew M Dai, Jonas Kemp, and Luke Metz · 2019
Earlier work this paper cites.
How does learning rate decay help modern neural networks?
Kaichao You, Mingsheng Long, Jianmin Wang, and Michael I Jordan · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
PIQA: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Earlier work this paper cites.
Stochastic gradient and Langevin processes
Xiang Cheng, Dong Yin, Peter Bartlett, and Michael Jordan · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, et al · 2020
Earlier work this paper cites.
An exponential learning rate schedule for deep learning
Zhiyuan Li and Sanjeev Arora · 2020
Earlier work this paper cites.
Language models are few-shot learners
Ben Mann, N Ryder, M Subbiah, J Kaplan, P Dhariwal, A Neelakantan, P Shyam, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, et al · 2020
Earlier work this paper cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Earlier work this paper cites.
Investigating prior knowledge for challenging Chinese machine reading comprehension
Kai Sun, Dian Yu, Dong Yu, and Claire Cardie · 2020
Cited alongside, same era.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Cited alongside, same era.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborova · 2021
Cited alongside, same era.
Continuous vs. discrete optimization of deep neural networks
Omer Elkabetz and Nadav Cohen · 2021
Cited alongside, same era.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2024
Later among the works it cites.
Deepseek LLM: Scaling open-source language models with longtermism
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, et al · 2024
Later among the works it cites.
A dynamical model of neural scaling laws
Blake Bordelon, Alexander Atanasov, and Cengiz Pehlevan · 2024
Later among the works it cites.
The road less scheduled
Aaron Defazio, Xingyu Yang, Harsh Mehta, Konstantin Mishchenko, Ahmed Khaled, and Ashok Cutkosky · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tatsunori Hashimoto · 2021
Cited alongside, same era.
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish · 2021
Cited alongside, same era.
Marcus Hutter · 2021
Cited alongside, same era.
Autodrop: Training deep learning models with automatic learning rate drop
Yunfei Teng, Jing Wang, and Anna Choromanska · 2021
Cited alongside, same era.
Revisiting neural scaling laws in language and vision
Ibrahim M Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai · 2022
Cited alongside, same era.
Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, et al · 2022
Cited alongside, same era.
Ce Ge, Zhijian Ma, Daoyuan Chen, Yaliang Li, and Bolin Ding · 2024
Later among the works it cites.
Scaling laws for data filtering– data curation cannot be compute agnostic
Sachin Goyal, Pratyush Maini, Zachary C. Lipton, Aditi Raghunathan, and J. Zico Kolter · 2024
Later among the works it cites.
OLMo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, et al · 2024
Later among the works it cites.
Scaling laws and compute-optimal training beyond fixed training durations
Alexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro Von Werra, and Martin Jaggi · 2024
Later among the works it cites.
MiniCPM: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Ganqu Cui, Chaoqun He, Weilin Zhao, Xiang Long, et al · 2024
Later among the works it cites.
Simple and scalable strategies to continually pre-train large language models
Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats Leon Richter, Quentin Gregory Anthony, Eugene Belilovsky, Timothée Lesort, et al · 2024
Later among the works it cites.
Scaling laws for learning with real and surrogate data
Ayush Jain, Andrea Montanari, and Eren Sasoglu · 2024
Later among the works it cites.
Scaling laws in linear regression: Compute, parameters, and data
Licong Lin, Jingfeng Wu, Sham M. Kakade, Peter L. Bartlett, and Jason D. Lee · 2024
Later among the works it cites.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, et al · 2024
Later among the works it cites.
Scaling laws for fine-grained mixture of experts
Jan Ludziejewski, Jakub Krajewski, Kamil Adamczewski, Maciej Pióro, Michał Krutul, Szymon Antoniak, Kamil Ciebiera, et al · 2024
Later among the works it cites.
An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem
Yoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee, Chris Mingard, and Ard A. Louis · 2024
Later among the works it cites.
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, et al · 2024
Later among the works it cites.
4+3 phases of compute-optimal neural scaling laws
Elliot Paquette, Courtney Paquette, Lechao Xiao, and Jeffrey Pennington · 2024
Later among the works it cites.
Power scheduler: A batch size and token number agnostic learning rate scheduler
Yikang Shen, Matthew Stallone, Mayank Mishra, Gaoyuan Zhang, Shawn Tan, Aditya Prasad, Adriana Meza Soria, et al · 2024
Later among the works it cites.
Scaling law with learning rate annealing
Howe Tissue, Venus Wang, and Lu Wang · 2024
Later among the works it cites.
Optimization hyper-parameter laws for large language models
Xingyu Xie, Kuangyu Ding, Shuicheng Yan, Kim-Chuan Toh, and Tianwen Wei · 2024
Later among the works it cites.
Data mixing laws: Optimizing data mixtures by predicting language modeling performance
Jiasheng Ye, Peiju Liu, Tianxiang Sun, Yunhua Zhou, Jun Zhan, and Xipeng Qiu · 2024
Later among the works it cites.
Straight to zero: Why linearly decaying the learning rate to zero works best for LLMs
Shane Bergsma, Nolan Simran Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, and Joel Hestness · 2025
Closest in time.
How feature learning can improve neural scaling laws
Blake Bordelon, Alexander Atanasov, and Cengiz Pehlevan · 2025
Closest in time.
Loss-to-loss prediction: Scaling laws for all datasets
David Brandfonbrener, Nikhil Anand, Nikhil Vyas, Eran Malach, and Sham M. Kakade · 2025
Closest in time.
Understanding optimization in deep learning with central flows
Jeremy Cohen, Alex Damian, Ameet Talwalkar, J Zico Kolter, and Jason D. Lee · 2025
Closest in time.
A solvable attention for neural scaling laws
Bochen Lyu, Di Wang, and Zhanxing Zhu · 2025
Closest in time.
Fabian Schaipp, Alexander Hägele, Adrien Taylor, Umut Simsekli, and Francis Bach · 2025
Closest in time.
Understanding warmup-stable-decay learning rates: A river valley loss landscape view
Kaiyue Wen, Zhiyuan Li, Jason S. Wang, David Leo Wright Hall, Percy Liang, and Tengyu Ma · 2025
Closest in time.