Fetching the paper…
Reading the bibliography…
Deep learning has arguably achieved tremendous success in recent years.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Jack Kiefer, Jacob Wolfowitz, et al · 1952
Earlier work this paper cites.
Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex
David H Hubel and Torsten N Wiesel · 1962
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
On the structure of continuous functions of several variables
David A Sprecher · 1965
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
VN Vapnik and A Ya Chervonenkis · 1971
Earlier work this paper cites.
A finite sample distribution-free performance bound for local discrimination rules
William H Rogers and Terry J Wagner · 1978
Earlier work this paper cites.
Distribution-free performance bounds for potential function rules
Luc Devroye and Terry Wagner · 1979
Earlier work this paper cites.
Adaptive estimation algorithms: convergence, optimality, stability
Boris Teodorovich Polyak and Yakov Zalmanovich Tsypkin · 1979
Earlier work this paper cites.
Projection pursuit regression
Jerome H Friedman and Werner Stuetzle · 1981
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
Kunihiko Fukushima and Sei Miyake · 1982
Earlier work this paper cites.
Optimal global rates of convergence for nonparametric regression
Charles J Stone · 1982
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1985
Earlier work this paper cites.
Sliced inverse regression for dimension reduction
Ker-Chau Li · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Ideal spatial adaptation by wavelet shrinkage
David L Donoho and Jain M Johnstone · 1994
Earlier work this paper cites.
Circuit complexity and neural networks
Ian Parberry · 1994
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
Heuristics of instability and stabilization in model selection
Leo Breiman et al · 1996
Earlier work this paper cites.
Random approximants and neural networks
Yuly Makovoz · 1996
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
Hrushikesh N Mhaskar · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
Online learning and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Density estimation for statistics and data analysis
Bernard W Silverman · 1998
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
High-dimensional data analysis: The curses and blessings of dimensionality
David L Donoho · 2000
Earlier work this paper cites.
On the near optimality of the stochastic approximation of smooth functions by neural networks
VE Maiorov and Ron Meir · 2000
Earlier work this paper cites.
Variable selection via nonconcave penalized likelihood and its oracle properties
Jianqing Fan and Runze Li · 2001
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
Harold Kushner and G George Yin · 2003
Earlier work this paper cites.
Fisher lecture: Dimension reduction in regression
R Dennis Cook et al · 2007
Earlier work this paper cites.
Efficient learning of sparse representations with an energy-based model
Christopher Poultney, Sumit Chopra, Yann LeCun, et al · 2007
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Earlier work this paper cites.
On functions of three variables
Vladimir I Arnold · 2009
Earlier work this paper cites.
Computational complexity: a modern approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
The power of convex relaxation: Near-optimal matrix completion
Emmanuel J Candès and Terence Tao · 2009
Earlier work this paper cites.
Deep boltzmann machines
Ruslan Salakhutdinov and Geoffrey Hinton · 2009
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Matrix computations
Gene H Golub and Charles F Van Loan · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Min Lin, Qiang Chen, and Shuicheng Yan · 2013
Cited alongside, same era.
Why does deep and cheap learning work so well?
Henry W Lin, Max Tegmark, and David Rolnick · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review
Tomaso Poggio, Hrushikesh Mhaskar, Lorenzo Rosasco, Brando Miranda, and Qianli Liao · 2017
Later among the works it cites.
The power of deeper networks for expressing natural functions
David Rolnick and Max Tegmark · 2017
Later among the works it cites.
Nonparametric regression using deep neural networks with relu activation function
Johannes Schmidt-Hieber · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Cited alongside, same era.
Dropout training as adaptive regularization
Stefan Wager, Sida Wang, and Percy S Liang · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Later among the works it cites.
Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations
E Weinan, Jiequn Han, and Arnulf Jentzen · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Later among the works it cites.
The deeptune framework for modeling and characterizing neurons in visual cortex area v4
Reza Abbasi-Asl, Yuansi Chen, Adam Bloniarz, Michael Oliver, Ben DB Willmore, Jack L Gallant, and Bin Yu · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Later among the works it cites.
Approximability of discriminators implies diversity in GANs
Yu Bai, Tengyu Ma, and Andrej Risteski · 2018
Later among the works it cites.
Deep learning and its applications in biomedicine
Chensi Cao, Feng Liu, Hai Tan, Deshou Song, Wenjie Shu, Weizhong Li, Yiming Zhou, Xiaochen Bo, and Zhi Xie · 2018
Later among the works it cites.
Neural ordinary differential equations
Tianqi Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud · 2018
Later among the works it cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Later among the works it cites.
Clinically applicable deep learning for diagnosis and referral in retinal disease
Jeffrey De Fauw, Joseph R Ledsam, Bernardino Romera-Paredes, Stanislav Nikolov, Nenad Tomasev, Sam Blackwell, Harry Askham, Xavier Glorot, Brendan O’Donoghue, Daniel Visentin, et al · 2018
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Later among the works it cites.
Local geometry of one-hidden-layer neural networks for logistic regression
Haoyu Fu, Yuejie Chi, and Yingbin Liang · 2018
Later among the works it cites.
Robust estimation and generative adversarial nets
Chao Gao, Jiyi Liu, Yuan Yao, and Weizhi Zhu · 2018
Later among the works it cites.
Learning one convolutional layer with overlapping patches
Surbhi Goel, Adam Klivans, and Raghu Meka · 2018
Later among the works it cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Later among the works it cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
On tighter generalization bound for deep neural networks: Cnns, resnets, and beyond
Xingguo Li, Junwei Lu, Zhaoran Wang, Jarvis Haupt, and Tuo Zhao · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
On the connection between learning two-layers neural networks and tensor decomposition
Marco Mondelli and Andrea Montanari · 2018
Later among the works it cites.
Do cifar-10 classifiers generalize to cifar-10?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Yaniv Romano, Matteo Sesia, and Emmanuel J Candès · 2018
Later among the works it cites.
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Later among the works it cites.
Hierarchical interpretations for neural network predictions
Chandan Singh, W James Murdoch, and Bin Yu · 2018
Later among the works it cites.
Mean field analysis of neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Later among the works it cites.
Can SGD Learn Recurrent Neural Networks with Provable Generalization?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Closest in time.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Closest in time.
Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval
Yuxin Chen, Yuejie Chi, Jianqing Fan, and Cong Ma · 2019
Closest in time.
Yuxin Chen, Yuejie Chi, Jianqing Fan, Cong Ma, and Yuling Yan · 2019
Closest in time.
A priori estimates of the population risk for residual networks
Weinan E, Chao Ma, and Qingcan Wang · 2019
Closest in time.
High probability generalization bounds for uniformly stable algorithms with nearly optimal rate
Vitaly Feldman and Jan Vondrak · 2019
Closest in time.
Analysis of a two-layer neural network via displacement convexity
Adel Javanmard, Marco Mondelli, and Andrea Montanari · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.