Fetching the paper…
Reading the bibliography…
To improve performance in contemporary deep learning, one is interested in scaling up the neural network in terms of both the number and the size of the layers.
Bayesian Learning for Neural Networks
Radford M. Neal · 1994
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Path-SGD: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan Salakhutdinov and Nathan Srebro · 2015
Earlier work this paper cites.
Fast DNN training based on auxiliary function technique
Dung T. Tran, Nobutaka Ono and Emmanuel Vincent · 2015
Earlier work this paper cites.
Theano: A Python framework for fast computation of mathematical expressions
Rami Al-Rfou, Guillaume Alain, Amjad Almahairi, Christof Angermueller, Dzmitry Bahdanau et al · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
On the convergence of Adam and beyond
Sashank J. Reddi, Satyen Kale and Sanjiv Kumar · 2018
Earlier work this paper cites.
Deep neural networks as Gaussian processes
Jaehoon Lee, Jascha Sohl-Dickstein, Jeffrey Pennington, Roman Novak, Sam Schoenholz et al · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel and Clement Hongler · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary et al · 2018
Earlier work this paper cites.
The spelled-out intro to neural networks and backpropagation: Building micrograd, 2018
Andrej Karpathy · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama and Yuichi Yoshida · 2018
Earlier work this paper cites.
Measuring and regularizing networks in function space
Ari Benjamin, David Rolnick and Konrad Kording · 2019
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei et al · 2019
Cited alongside, same era.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Cited alongside, same era.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra and Ali Jadbabaie · 2020
Cited alongside, same era.
On the distance between two neural networks and the stability of learning
Jeremy Bernstein, Arash Vahdat, Yisong Yue and Ming-Yu Liu · 2020
Cited alongside, same era.
On the linearity of large non-linear models: When and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu and Mikhail Belkin · 2020
Depth dependence of μ \mu P learning rates in ReLU MLPs
Samy Jelassi, Boris Hanin, Ziwei Ji, Sashank J. Reddi, Srinadh Bhojanapalli et al · 2023
Later among the works it cites.
Topics in convex optimisation
Hamza Fawzi · 2023
Later among the works it cites.
Convex and non-convex optimization under generalized smoothness
Haochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin and Ali Jadbabaie · 2023
Later among the works it cites.
A spectral condition for feature learning
Greg Yang, James B. Simon and Jeremy Bernstein · 2023
Later among the works it cites.
Efficient parametric approximations of neural network function space distance
Nikita Dhawan, Sicong Huang, Juhan Bae and Roger Grosse · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu et al · 2021
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter and Ameet Talwalkar · 2021
Cited alongside, same era.
Tensor programs IV: Feature learning in infinite-width neural networks
Greg Yang and J. Edward Hu · 2021
Cited alongside, same era.
The large learning rate phase of deep learning, 2021
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein and Guy Gur-Ari · 2021
Cited alongside, same era.
The neural covariance SDE: Shaped infinite depth-and-width networks at initialization
Mufan Bill Li, Mihai Nica and Daniel M. Roy · 2022
Cited alongside, same era.
The Principles of Deep Learning Theory
Daniel A. Roberts, Sho Yaida and Boris Hanin · 2022
Cited alongside, same era.
Namhoon Cho and Hyo-Sang Shin · 2023
Later among the works it cites.
Automatic gradient descent: Deep learning without hyperparameters
Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan and Yisong Yue · 2023
Later among the works it cites.
Learning-rate-free learning by D-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Later among the works it cites.
DoG is SGD’s best friend: A parameter-free dynamic step size schedule
Maor Ivgi, Oliver Hinder and Yair Carmon · 2023
Later among the works it cites.
TinyStories: How small can language models be and still speak coherent English?
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
Tensor programs VI: Feature learning in infinite depth neural networks
Greg Yang, Dingli Yu, Chen Zhu and Soufiane Hayou · 2024
Closest in time.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
Blake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin and Cengiz Pehlevan · 2024
Closest in time.
A large-scale exploration of μ \mu -transfer
Lucas Lingle · 2024
Closest in time.
Prodigy: An expeditiously adaptive parameter-free learner
Konstantin Mishchenko and Aaron Defazio · 2024
Closest in time.