Fetching the paper…
Reading the bibliography…
The push to train ever larger neural networks has motivated the study of initialization and training at large network width.
Problem complexity and method efficiency in optimization
Arkady S. Nemirovsky and David B. Yudin · 1983
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Efficient backprop
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2002
Earlier work this paper cites.
Matrix nearness problems with Bregman divergences
Inderjit S. Dhillon and Joel A. Tropp · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Non-asymptotic theory of random matrices: Extreme singular values
Mark Rudelson and Roman Vershynin · 2010
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Scaling SGD batch size to 32K for ImageNet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
signSGD: Compressed Optimisation for Non-Convex Problems, February 2018
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Clément Hongler, and Franck Gabriel · 2018
Cited alongside, same era.
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Spectral normalization for Generative Adversarial Networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
On the infinite width limit of neural networks with a standard parameterization
Jascha Sohl-Dickstein, Roman Novak, Samuel S. Schoenholz, and Jaehoon Lee · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Later among the works it cites.
Learning by turning: Neural architecture aware optimisation
Yang Liu, Jeremy Bernstein, Markus Meister, and Yisong Yue · 2021
Later among the works it cites.
Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roman Vershynin · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Layer rotation: A surprisingly simple indicator of generalization in deep networks?
Simon Carbonnelle and Christophe De Vleeschouwer · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Generalizable adversarial training via spectral normalization
Farzan Farnia, Jesse Zhang, and David Tse · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Tensor Programs V: Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
Later among the works it cites.
Neural networks as kernel learners: the silent alignment effect
Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan · 2022
Later among the works it cites.
Self-consistent dynamical field theory of kernel evolution in wide neural networks
Blake Bordelon and Cengiz Pehlevan · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
Grant M. Rotskoff and Eric Vanden-Eijnden · 2022
Later among the works it cites.
Mean field analysis of deep neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2022
Later among the works it cites.
Limitations of the NTK for understanding generalization in deep learning
Nikhil Vyas, Yamini Bansal, and Preetum Nakkiran · 2022
Later among the works it cites.
Meta-principled family of hyperparameter scaling strategies
Sho Yaida · 2022
Later among the works it cites.
Automatic Gradient Descent: Deep Learning without Hyperparameters
Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan, and Yisong Yue · 2023
Closest in time.
Cerebras-GPT: Open compute-optimal language models trained on the Cerebras wafer-scale cluster
Nolan Dey, Gurpreet Gosal, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, Joel Hestness, et al · 2023
Closest in time.
Tensor programs IVb: Adaptive optimization in the ∞ \infty -width limit
Greg Yang and Etai Littwin · 2023
Closest in time.