Fetching the paper…
Reading the bibliography…
We introduce a new second-order inertial optimization method for machine learning called INNA.
An inertial Newton algorithm for deep learning
Castera, C., Bolte, J., Févotte, C., and Pauwels, E. (2019) · 1905
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T. (1964) · 1964
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
Ljung, L. (1977) · 1977
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E. and Hinton, G. E. (1986) · 1986
Earlier work this paper cites.
Optimization and nonsmooth analysis
Clarke, F. H. (1990) · 1990
Earlier work this paper cites.
Python reference manual
Rossum, G. (1995) · 1995
Earlier work this paper cites.
A dynamical system associated with Newton’s method for parametric approximations of convex minimization problems
Alvarez, F. and Pérez, J. M. (1998) · 1998
Earlier work this paper cites.
Tame topology and o-minimal structures
van den Dries, L. (1998) · 1998
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
Kurdyka, K. (1998) · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., Haffner, P., et al. (1998) · 1998
Earlier work this paper cites.
Dynamics of stochastic approximation algorithms
Benaïm, M. (1999) · 1999
Earlier work this paper cites.
An introduction to o-minimal geometry
Coste, M. (2000) · 2000
Earlier work this paper cites.
A second-order gradient-like dissipative dynamical system with Hessian-driven damping: Application to optimization and mechanics
Alvarez, F., Attouch, H., Bolte, J., and Redont, P. (2002) · 2002
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
Kushner, H. and Yin, G. G. (2003) · 2003
Earlier work this paper cites.
Stochastic approximations and differential inclusions
Benaïm, M., Hofbauer, J., and Sorin, S. (2005) · 2005
Earlier work this paper cites.
Convergence of constant step stochastic gradient descent for non-smooth non-convex functions
Bianchi, P., Hachem, W., and Schechtman, S. (2020) · 2005
Cited alongside, same era.
Matplotlib: A 2D graphics environment
Hunter, J. D. (2007) · 2007
Cited alongside, same era.
The tradeoffs of large scale learning
Bottou, L. and Bousquet, O. (2008) · 2008
Cited alongside, same era.
Stochastic approximation: A dynamical systems viewpoint
Borkar, V. S. (2009) · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Cited alongside, same era.
Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the Kurdyka-Łojasiewicz inequality
Attouch, H., Bolte, J., Redont, P., and Soubeyran, A. (2010) · 2010
Network in Network
Lin, M., Chen, Q., and Yan, S. (2014) · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Later among the works it cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al. (2016) · 2016
Later among the works it cites.
A stochastic quasi-Newton method for large-scale optimization
Byrd, R. H., Hansen, S. L., Nocedal, J., and Singer, Y. (2016) · 2016
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B. (2017) · 2017
Later among the works it cites.
Opérateurs monotones aléatoires et application à l’optimisation stochastique
Adil, S. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Characterizations of Łojasiewicz inequalities: subgradient flows, talweg, convexity
Bolte, J., Daniilidis, A., Ley, O., and Mazet, L. (2010) · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
On the use of stochastic Hessian information in optimization methods for machine learning
Byrd, R. H., Chin, G. M., Neveitt, W., and Nocedal, J. (2011) · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and Bach, F. R. (2011) · 2011
Cited alongside, same era.
The NumPy array: a structure for efficient numerical computation
Walt, S. v. d., Colbert, S. C., and Varoquaux, G. (2011) · 2011
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J. (2018) · 2018
Later among the works it cites.
Stochastic methods for composite and weakly convex optimization problems
Duchi, J. C. and Ruan, F. (2018) · 2018
Later among the works it cites.
INNA for deep learning
Castera, C. (2019) · 2019
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Closest in time.
First-order optimization algorithms via inertial systems with Hessian driven damping
Attouch, H., Chbani, Z., Fadili, J., and Riahi, H. (2020) · 2020
Closest in time.
An investigation of Newton-sketch and subsampled Newton methods
Berahas, A. S., Bollapragada, R., and Nocedal, J. (2020) · 2020
Closest in time.
Stochastic subgradient method converges on tame functions
Davis, D., Drusvyatskiy, D., Kakade, S., and Lee, J. D. (2020) · 2020
Closest in time.
Second-order optimization for non-convex machine learning: an empirical study
Xu, P., Roosta, F., and Mahoney, M. W. (2020) · 2020
Closest in time.
Convergence and dynamical behavior of the ADAM algorithm for nonconvex stochastic optimization
Barakat, A. and Bianchi, P. (2021) · 2021
Closest in time.