Fetching the paper…
Reading the bibliography…
Large Language Models have driven significant AI advancements, yet their training is resource-intensive and highly sensitive to hyper-parameter selection.
Brownian motion in a field of force and the diffusion model of chemical reactions
H. A. Kramers · 1940
Earlier work this paper cites.
Stochastic Differential Equations
I. I. Gihman and A. V. Skorohod · 1979
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Introduction to Optimization
B. Polyak · 1987
Earlier work this paper cites.
Numerical Solution to Stochastic Differential Equations
P. E. Kloeden and E. Platen · 1999
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
D. P. Bertsekas and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Distributional and L q L^{q} norm inequalities for polynomials over convex bodies in ℝ n \mathbb{R}^{n}
A. Carbery and J. Wright · 2001
Earlier work this paper cites.
Gaussian process approximations of stochastic differential equations
C. Archambeau, D. Cornford, M. Opper, and J. Shawe-Taylor · 2007
Earlier work this paper cites.
Spectral analysis of large dimensional random matrices , volume 20
Z. Bai and J. W. Silverstein · 2010
Earlier work this paper cites.
Small deviations for beta ensembles
M. Ledoux and B. Rider · 2010
Earlier work this paper cites.
Kramers’ law: Validity, derivations and generalisations
N. Berglund · 2013
Earlier work this paper cites.
Sharp estimates for metastable lifetimes in parabolic SPDEs: Kramers’ law and beyond
N. Berglund and B. Gentz · 2013
Earlier work this paper cites.
Gaussian filtering and smoothing for continuous-discrete dynamic systems
S. Särkkä and J. Sarmavuori · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
A differential equation for modeling Nesterov’s accelerated gradient method: Theory and insights
W. Su, S. Boyd, and E. Candes · 2014
Earlier work this paper cites.
Metastability: a potential-theoretic approach
A. Bovier and F. Den Hollander · 2015
Earlier work this paper cites.
Posterior inference on parameters of stochastic differential equations via non-linear gaussian filtering and adaptive MCMC
S. Särkkä, J. Hartikainen, I. S. Mbalawata, and H. Haario · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Earlier work this paper cites.
Three factors influencing minima in SGD
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, et al · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Q. Li, C. Tai, and E. Weinan · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Asymptotic for a second-order evolution equation with convex potential andvanishing damping term
R. May · 2017
Earlier work this paper cites.
Fast convergence of inertial dynamics and algorithms with asymptotic vanishing viscosity
H. Attouch, Z. Chbani, J. Peypouquet, and P. Redont · 2018
Earlier work this paper cites.
Stochastic methods for composite and weakly convex optimization problems
J. C. Duchi and F. Ruan · 2018
Earlier work this paper cites.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Earlier work this paper cites.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
F. He, T. Liu, and D. Tao · 2019
Earlier work this paper cites.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Q. Li, C. Tai, and E. Weinan · 2019
Earlier work this paper cites.
A dynamical systems perspective on Nesterov acceleration
M. Muehlebach and M. Jordan · 2019
Earlier work this paper cites.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
T. H. Nguyen, U. Simsekli, M. Gurbuzbalaban, and G. Richard · 2019
Earlier work this paper cites.
Applied Stochastic Differential Equations , volume 10
S. Särkkä and A. Solin · 2019
Earlier work this paper cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
U. Simsekli, L. Sagun, and M. Gurbuzbalaban · 2019
Earlier work this paper cites.
Z. Zhu, J. Wu, B. Yu, L. Wu, and J. Ma · 2019
Cited alongside, same era.
Stochastic subgradient method converges on tame functions
D. Davis, D. Drusvyatskiy, S. Kakade, and J. D. Lee · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, et al · 2020
Cited alongside, same era.
A constructive prediction of the generalization error across scales
J. S. Rosenfeld, A. Rosenfeld, Y. Belinkov, and N. Shavit · 2020
Continual pre-training of language models
Z. Ke, Y. Shao, H. Lin, T. Konishi, G. Kim, and B. Liu · 2023
Later among the works it cites.
Full parameter fine-tuning for large language models with limited resources
K. Lv, Y. Yang, T. Liu, Q. Gao, Q. Guo, and X. Qiu · 2023
Later among the works it cites.
Repofusion: Training code models to understand your repository
D. Shrivastava, D. Kocetkov, H. de Vries, D. Bahdanau, and T. Scholak · 2023
Later among the works it cites.
An elementary proof of anti-concentration for degree two non-negative gaussian polynomials
S. Tu and R. Boczar · 2023
Later among the works it cites.
Convergence guarantees for stochastic subgradient methods in nonsmooth nonconvex optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Z. Xie, I. Sato, and M. Sugiyama · 2020
Cited alongside, same era.
A novel convergence analysis for algorithms of the Adam family
Z. Guo, Y. Xu, W. Yin, R. Jin, and T. Yang · 2021
Cited alongside, same era.
D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish · 2021
Cited alongside, same era.
The Lipschitz constant of self-attention
H. Kim, G. Papamakarios, and A. Mnih · 2021
Cited alongside, same era.
On the validity of modeling SGD with stochastic differential equations (SDEs)
Z. Li, S. Malladi, and S. Arora · 2021
Cited alongside, same era.
Scalable inference in SDEs by direct matching of the Fokker–Planck–Kolmogorov equation
A. Solin, E. Tamir, and P. Verma · 2021
Cited alongside, same era.
Shape matters: Understanding the implicit bias of the noise covariance
H. Zhang, C. Wei, J. Lee, and T. Ma · 2021
Cited alongside, same era.
N. Xiao, X. Hu, and K.-C. Toh · 2023
Later among the works it cites.
Fast convex optimization via a third-order in time evolution equation: Toges-v an improved version of toges
H. Attouch, Z. Chbani, and H. Riahi · 2024
Closest in time.
Revisiting the noise model of stochastic gradient descent
B. Battash, L. Wolf, and O. Lindenbaum · 2024
Closest in time.
Chinchilla Scaling: A replication attempt
T. Besiroglu, E. Erdil, M. Barnett, and J. You · 2024
Closest in time.
DeepSeek LLM: Scaling open-source language models with longtermism
X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, et al · 2024
Closest in time.
Stochastic differential equations for modeling first order optimization methods
M. Dambrine, C. Dossal, B. Puig, and A. Rondepierre · 2024
Closest in time.
DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model
DeepSeek-AI, A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, et al · 2024
Closest in time.
Stochastic Bregman Subgradient Methods for Nonsmooth Nonconvex Optimization Problems
K. Ding and K.-C. Toh · 2024
Closest in time.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, et al · 2024
Closest in time.
Scaling laws for data filtering–data curation cannot be compute agnostic
S. Goyal, P. Maini, Z. C. Lipton, A. Raghunathan, and J. Z. Kolter · 2024
Closest in time.
Accelerated objective gap and gradient norm convergence for gradient descent via long steps
B. Grimmer, K. Shu, and A. Wang · 2024
Closest in time.
Efficient continual pre-training by mitigating the stability gap
Y. Guo, J. Fu, H. Zhang, D. Zhao, and Y. Shen · 2024
Closest in time.
Scaling laws and compute-optimal training beyond fixed training durations
A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. Von Werra, and M. Jaggi · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, et al · 2024
Closest in time.
Simple and scalable strategies to continually pre-train large language models
A. Ibrahim, B. Thérien, K. Gupta, M. L. Richter, Q. Anthony, T. Lesort, et al · 2024
Closest in time.
Scaling laws for downstream task performance of large language models
B. Isik, N. Ponomareva, H. Hazimeh, D. Paparas, S. Vassilvitskii, and S. Koyejo · 2024
Closest in time.
Some fundamental aspects about Lipschitz continuity of neural networks
G. Khromov and S. P. Singh · 2024
Closest in time.
R. Maulen-Soto, J. Fadili, H. Attouch, and P. Ochs · 2024
Closest in time.
Scaling data-constrained language models
N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, et al · 2024
Closest in time.
Reuse, don’t retrain: A recipe for continued pretraining of language models
J. Parmar, S. Satheesh, M. Patwary, M. Shoeybi, and B. Catanzaro · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, et al · 2024
Closest in time.
T. Rotaru, F. Glineur, and P. Patrinos · 2024
Closest in time.
Beyond Chinchilla-optimal: Accounting for inference in language model scaling laws
N. Sardana, J. Portes, S. Doubov, and J. Frankle · 2024
Closest in time.
Skywork-MoE: A deep dive into training techniques for mixture-of-experts language models
T. Wei, B. Zhu, L. Zhao, C. Cheng, B. Li, W. Lü, et al · 2024
Closest in time.
Adam-family methods for nonsmooth optimization with convergence guarantees
N. Xiao, X. Hu, X. Liu, and K.-C. Toh · 2024
Closest in time.
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, et al · 2024
Closest in time.
LongSkywork: A training recipe for efficiently extending context length in large language models
L. Zhao, T. Wei, L. Zeng, C. Cheng, L. Yang, P. Cheng, et al · 2024
Closest in time.