Fetching the paper…
Reading the bibliography…
Deep learning has aroused extensive attention due to its great empirical success.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Principles of neurodynamics: perceptrons and the theory of brain mechanisms
Rosenblatt, F · 1961
Earlier work this paper cites.
Une propriété topologique des sous-ensembles analytiques réels. In: Les Équations aux dérivées partielles
Łojasiewicz, S · 1963
Earlier work this paper cites.
Ensembles semi-analytiques
Łojasiewicz, S · 1965
Earlier work this paper cites.
Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K · 1980
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Sur la geometrie semi-et sous-analytique
Łojasiewicz, S · 1993
Earlier work this paper cites.
Geometry of subanalytic and semialgebraic sets , volume 150 of Progress in Mathematics
Shiota, M · 1997
Earlier work this paper cites.
Real algebraic geometry , volume 3
Bochnak, J., Coste, M., and Roy, M.-F · 1998
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
Kurdyka, K · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Variational analysis
Rockafellar, R. T. and Wets, R. J.-B · 1998
Earlier work this paper cites.
A primer of real analytic functions
Krantz, S. and Parks, H. R · 2002
Earlier work this paper cites.
Regularization and variable selection via the elastic net
Zou, H. and Hastie, T · 2005
Earlier work this paper cites.
Variational analysis and generalized differentiation I: Basic Theory
Mordukhovich, B. S · 2006
Earlier work this paper cites.
On the convergence of the proximal algorithm for nonsmooth functions involving analytic features
Attouch, H. and Bolte, J · 2009
Earlier work this paper cites.
Proximal alternating minimization and projection methods for nonconvex problems: an approach based on the Kurdyka-Łojasiewicz inequality
Attouch, H., Bolte, J., Redont, P., and Soubeyran, A · 2010
Cited alongside, same era.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Boyd, S., Parikh, N., Chu, E., Peleato, B., Eckstein, J., et al · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
On optimization methods for deep learning
Le, Q. V., Ngiam, J., Coates, A., Lahiri, A., Prochnow, B., and Ng, A. Y · 2011
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G., Deng, L., Yu, D., Dahl, G. E., Mohamed, A.-r., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., and Kingsbury, B · 2012
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Later among the works it cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Later among the works it cites.
Training neural networks without gradients: A scalable ADMM approach
Taylor, G., Burmeister, R., Xu, Z., Singh, B., Patel, A., and Goldstein, T · 2016
Later among the works it cites.
Efficient training of very deep neural networks for supervised hashing
Zhang, Z., Chen, Y., and Saligrama, V · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Cited alongside, same era.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Cited alongside, same era.
Lecture 6.5-RMSProp: Divide the gradient by running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward-backward splitting, and regularized Gauss-Seidel methods
Attouch, H., Bolte, J., and Svaiter, B. F · 2013
Cited alongside, same era.
Deep convolutional neural networks for LVCSR
Sainath, T. N., Mohamed, A.-r., Kingsbury, B., and Ramabhadran, B · 2013
Cited alongside, same era.
A block coordinate descent method for regularized multiconvex optimization with applications to nonnegative tensor factorization and completion
Xu, Y. and Yin, W · 2013
Cited alongside, same era.
Proximal alternating linearized minimization for nonconvex and nonsmooth problems
Bolte, J., Sabach, S., and Teboulle, M · 2014
Cited alongside, same era.
Liao, Q. and Poggio, T · 2017
Later among the works it cites.
A distributed block coordinate descent method for training ℓ 1 \ell_{1} regularized linear classifiers
Mahajan, D., Keerthi, S. S., and Sundararajan, S · 2017
Later among the works it cites.
A globally convergent algorithm for nonconvex optimization based on block coordinate update
Xu, Y. and Yin, W · 2017
Later among the works it cites.
Convergent block coordinate descent for training Tikhonov regularized deep neural networks
Zhang, Z. and Brand, M · 2017
Later among the works it cites.
Askari, A., Negiar, G., Sambharya, R., and EI Ghaoui, L · 2018
Closest in time.
Fenchel lifted networks: a Lagrange relaxation of neural network training
Gu, F., Askari, A., and El Ghaoui, L · 2018
Closest in time.
A proximal block coordinate descent algorithm for deep neural network training
Lau, T. T.-K., Zeng, J., Wu, B., and Yao, Y · 2018
Closest in time.
Lectures on convex optimization
Nesterov, Y · 2018
Closest in time.
On the convergence of Adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Closest in time.
Stochastic subgradient method converges on tame functions
Davis, D., Drusvyatskiy, D., Kakade, S., and Lee, J. D · 2019
Closest in time.
Global convergence of admm in nonconvex nonsmooth optimization
Wang, Y., Yin, W., and Zeng, J · 2019
Closest in time.