Fetching the paper…
Reading the bibliography…
Beyond neural scaling laws, little is known about the laws underlying large language models (LLMs).
A learning algorithm for boltzmann machines
David H Ackley, Geoffrey E Hinton, and Terrence J Sejnowski · 1985
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel · 2001
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Mandt Stephan, Matthew D Hoffman, David M Blei, et al · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Earlier work this paper cites.
How noise affects the hessian spectrum in overparameterized neural networks
Mingwei Wei and David J Schwab · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Cited alongside, same era.
Statistical mechanics of deep learning
Yasaman Bahri, Jonathan Kadmon, Jeffrey Pennington, Sam S Schoenholz, Jascha Sohl-Dickstein, and Surya Ganguli · 2020
Cited alongside, same era.
Implicit gradient regularization
David GT Barrett and Benoit Dherin · 2020
Cited alongside, same era.
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
The implicit bias for adaptive optimization algorithms on homogeneous neural networks
Bohan Wang, Qi Meng, Wei Chen, and Tie-Yan Liu · 2021
Later among the works it cites.
The limiting dynamics of sgd: Modified loss, phase-space oscillations, and anomalous diffusion
Daniel Kunin, Javier Sagastuy-Brena, Lauren Gillespie, Eshed Margalit, Hidenori Tanaka, Surya Ganguli, and Daniel LK Yamins · 2023
Later among the works it cites.
Stochastic collapse: How gradient noise attracts sgd dynamics towards simpler subnetworks
Feng Chen, Daniel Kunin, Atsushi Yamamura, and Surya Ganguli · 2023
Later among the works it cites.
Thermodynamics-inspired explanations of artificial intelligence
Shams Mehdi and Pratyush Tiwary · 2024
Later among the works it cites.
Understanding warmup-stable-decay learning rates: A river valley loss landscape perspective
Kaiyue Wen, Zhiyuan Li, Jason Wang, David Hall, Percy Liang, and Tengyu Ma · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Direction matters: On the implicit bias of stochastic gradient descent with moderate learning rate
Jingfeng Wu, Difan Zou, Vladimir Braverman, and Quanquan Gu · 2020
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Cited alongside, same era.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, et al · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Cited alongside, same era.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al · 2024
Later among the works it cites.
Scaling laws and compute-optimal training beyond fixed training durations
Alex Hägele, Elie Bakouch, Atli Kosson, Leandro Von Werra, Martin Jaggi, et al · 2024
Later among the works it cites.
modded-nanogpt
Jordan Keller · 2024
Later among the works it cites.
Focus: First order concentrated updating scheme
Yizhou Liu, Ziming Liu, and Jeff Gore · 2025
Closest in time.
A multi-power law for loss curve prediction across learning rate schedules
Kairong Luo, Haodong Wen, Shengding Hu, Zhenbo Sun, Maosong Sun, Zhiyuan Liu, Kaifeng Lyu, and Wenguang Chen · 2025
Closest in time.
An overview of condensation phenomenon in deep learning
Zhi-Qin John Xu, Yaoyu Zhang, and Zhangchen Zhou · 2025
Closest in time.