Fetching the paper…
Reading the bibliography…
The architecture of a deep neural network is defined explicitly in terms of the number of layers, the width of each layer and the general network topology.
Das asymptotische Verteilungsgesetz der Eigenwerte linearer partieller Differentialgleichungen (mit einer Anwendung auf die Theorie der Hohlraumstrahlung)
Hermann Weyl · 1912
Earlier work this paper cites.
Perturbation Theory for Linear Operators
Tosio Kato · 1966
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Lev M. Bregman · 1967
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkady S. Nemirovsky and David B. Yudin · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton and Ronald J. Williams · 1986
Earlier work this paper cites.
Numerical Methods for Least Squares Problems
Åke Björck · 1996
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 1999
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Yurii Nesterov and Boris Polyak · 2006
Earlier work this paper cites.
Perturbation of the SVD in the presence of small singular values
Michael Stewart · 2006
Earlier work this paper cites.
Matrix nearness problems with Bregman divergences
Inderjit S. Dhillon and Joel A. Tropp · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Multiplicative dynamics underlie the emergence of the log-normal distribution of spine sizes in the neocortex in vivo
Yonatan Loewenstein, Annerose Kuras and Simon Rumpel · 2011
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, X. Zhang, Shaoqing Ren and Jian Sun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen Zhu, Brian Bullins, Elad Hazan and Tengyu Ma · 2016
Cited alongside, same era.
Theano: A Python framework for fast computation of mathematical expressions
Rami Al-Rfou, Guillaume Alain, Amjad Almahairi, Christof Angermueller, Dzmitry Bahdanau, Nicolas Ballas, Frédéric Bastien, Justin Bayer, Anatoly Belikov, Alexander Belopolsky et al · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio and Aaron Courville · 2016
Cited alongside, same era.
MM Optimization Algorithms
Kenneth Lange · 2016
Cited alongside, same era.
Second-order stochastic optimization for machine learning in linear time
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga et al · 2019
Later among the works it cites.
On the distance between two neural networks and the stability of learning
Jeremy Bernstein, Arash Vahdat, Yisong Yue and Ming-Yu Liu · 2020
Later among the works it cites.
Exploring the role of loss functions in multiclass classification
Ahmet Demirkaya, Jiasi Chen and Samet Oymak · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan and Samy Bengio · 2020
Later among the works it cites.
ICML 2020 tutorial on parameter-free online optimization, 2020
Francesco Orabona and Ashok Cutkosky · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Naman Agarwal, Brian Bullins and Elad Hazan · 2017
Cited alongside, same era.
Train CIFAR-10 with PyTorch
Kuang Liu · 2017
Cited alongside, same era.
Are GANs created equal? A large-scale study
Mario Lucic, Karol Kurach, Marcin Michalski, Sylvain Gelly and Olivier Bousquet · 2017
Cited alongside, same era.
The exploding gradient problem demystified
George Philipp, Dawn Xiaodong Song and Jaime G. Carbonell · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Cited alongside, same era.
Scaling SGD batch size to 32K for ImageNet training
Yang You, Igor Gitman and Boris Ginsburg · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis and Jorge Nocedal · 2018
Cited alongside, same era.
Or Sharir, Barak Peleg and Yoav Shoham · 2020
Later among the works it cites.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra and Ali Jadbabaie · 2020
Later among the works it cites.
High-performance large-scale image recognition without normalization
Andy Brock, Soham De, Samuel L. Smith and Karen Simonyan · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter and Ameet Talwalkar · 2021
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs. cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2021
Later among the works it cites.
Descending through a crowded valley—benchmarking deep learning optimizers
Robin M. Schmidt, Frank Schneider and Philipp Hennig · 2021
Later among the works it cites.
Tensor programs IV: Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Later among the works it cites.
Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen and Jianfeng Gao · 2021
Later among the works it cites.
Optimisation & Generalisation in Networks of Neurons
Jeremy Bernstein · 2022
Later among the works it cites.
Investigating generalization by controlling normalized margin
Alexander R. Farhang, Jeremy Bernstein, Kushal Tirumala, Yang Liu and Yisong Yue · 2022
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu and Mikhail Belkin · 2022
Later among the works it cites.
Mirror descent maximizes generalized margin and can be implemented efficiently
Haoyuan Sun, Kwangjun Ahn, Christos Thrampoulidis and Navid Azizan · 2022
Later among the works it cites.