Fetching the paper…
Reading the bibliography…
We consider the problem of minimizing a sum of $n$ functions over a convex parameter set $\mathcal{C} \subset \mathbb{R}^p$ where $n\gg p\gg 1$.
Herbert Robbins and Sutton Monro, A stochastic approximation method , Annals of mathematical statistics (1951)
1951
Earlier work this paper cites.
Charles G Broyden, The convergence of a class of double-rank minimization algorithms 2. the new algorithm , IMA Journal of Applied Mathematics 6
1970
Earlier work this paper cites.
Roger Fletcher, A new approach to variable metric algorithms , The computer journal 13
1970
Earlier work this paper cites.
Donald Goldfarb, A family of variable-metric methods derived by variational means , Mathematics of computation 24
1970
Earlier work this paper cites.
David F Shanno, Conditioning of quasi-newton methods for function minimization , Mathematics of computation 24
1970
Earlier work this paper cites.
John E Dennis, Jr and Jorge J Moré, Quasi-newton methods, motivation and theory , SIAM review 19
1977
Earlier work this paper cites.
Jian-Feng Cai, Emmanuel J Candès, and Zuowei Shen, A singular value thresholding algorithm for matrix completion , SIAM Journal on Optimization 20
1982
Earlier work this paper cites.
Yurii Nesterov, A method for unconstrained convex minimization problem with the rate of convergence o (1/k2) , Doklady AN SSSR, vol. 269, 1983, pp. 543–547
1983
Earlier work this paper cites.
Christopher M. Bishop, Neural networks for pattern recognition , Oxford University Press, Inc., NY, USA, 1995
1995
Earlier work this paper cites.
Aad W Van der Vaart and Jon A Wellner, Weak convergence , Springer, 1996
1996
Earlier work this paper cites.
Shun-Ichi Amari, Natural gradient works efficiently in learning , Neural computation 10
1998
Earlier work this paper cites.
Vladimir Vapnik, Statistical learning theory , vol. 2, Wiley New York, 1998
1998
Earlier work this paper cites.
Jock A Blackard and Denis J Dean, Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables , Computers and electronics in agriculture 24
1999
Earlier work this paper cites.
Bernhard Schölkopf and Alexander J Smola, Learning with kernels: support vector machines, regularization, optimization, and beyond , MIT press, 2002
2002
Earlier work this paper cites.
Stephen Boyd and Lieven Vandenberghe, Convex optimization , Cambridge University Press, New York, NY, USA, 2004
2004
Earlier work this paper cites.
S Sathiya Keerthi and Dennis DeCoste, A modified finite newton method for fast solution of large scale linear svms , Journal of Machine Learning Research, 2005, pp. 341–361
2005
Earlier work this paper cites.
Olivier Chapelle, Training a support vector machine in the primal , Neural Computation 19
2007
Cited alongside, same era.
Nicolas L. Roux, Pierre antoine Manzagol, and Yoshua Bengio, Topmoumoute online natural gradient algorithm , Advances in Neural Information Processing Systems 20, 2008, pp. 849–856
2008
Cited alongside, same era.
Igor Griva, Stephen G Nash, and Ariela Sofer, Linear and nonlinear optimization , Siam, 2009
2009
Cited alongside, same era.
Léon Bottou, Large-scale machine learning with stochastic gradient descent , Proceedings of COMPSTAT’2010, Springer, 2010, pp. 177–186
2010
Cited alongside, same era.
2010
Cited alongside, same era.
Joel A Tropp, User-friendly tail bounds for sums of random matrices , Foundations of Computational Mathematics 12
2012
Later among the works it cites.
Oriol Vinyals and Daniel Povey, Krylov Subspace Descent for Deep Learning , The 15th International Conference on Artificial Intelligence and Statistics - (AISTATS-12), 2012
2012
Later among the works it cites.
2013
Later among the works it cites.
Paramveer Dhillon, Yichao Lu, Dean P Foster, and Lyle Ungar, New subsampling algorithms for fast least squares regression , Advances in Neural Information Processing Systems 26, 2013, pp. 360–368
2013
Later among the works it cites.
M. Lichman, UCI machine learning repository , 2013
2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
James Martens, Deep learning via hessian-free optimization , Proceedings of the 27th International Conference on Machine Learning (ICML-10), 2010, pp. 735–742
2010
Cited alongside, same era.
Nicolas L Roux and Andrew W Fitzgibbon, A fast natural newton method , Proceedings of the 27th International Conference on Machine Learning (ICML-10), 2010, pp. 623–630
2010
Cited alongside, same era.
2010
Cited alongside, same era.
Richard H Byrd, Gillian M Chin, Will Neveitt, and Jorge Nocedal, On the use of stochastic hessian information in optimization methods for machine learning , SIAM Journal on Optimization 21
2011
Cited alongside, same era.
Thierry Bertin-Mahieux, Daniel P.W. Ellis, Brian Whitman, and Paul Lamere, The million song dataset , Proceedings of the 12th International Conference on Music Information Retrieval (ISMIR 2011), 2011
2011
Cited alongside, same era.
John Duchi, Elad Hazan, and Yoram Singer, Adaptive subgradient methods for online learning and stochastic optimization , Journal of Machine Learning Research 12
2011
Cited alongside, same era.
Franz Graf, Hans-Peter Kriegel, Matthias Schubert, Sebastian Pölsterl, and Alexander Cavallaro, 2d image registration in ct images using radial image descriptors , Medical Image Computing and Computer-Assisted Intervention–MICCAI 2011, Springer, 2011, pp. 607–614
2011
Cited alongside, same era.
Later among the works it cites.
2013
Later among the works it cites.
Alan Senior, Georg Heigold, Marc’Aurelio Ranzato, and Ke Yang, An empirical study of learning rates in deep neural networks for speech recognition , Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on, IEEE, 2013, pp. 6724–6728
2013
Later among the works it cites.
2013
Later among the works it cites.
2014
Later among the works it cites.
Matan Gavish and David L Donoho, Optimal shrinkage of singular values , arXiv:1405.7511 (2014)
2014
Later among the works it cites.
Lester Mackey, Michael I Jordan, Richard Y Chen, Brendan Farrell, Joel A Tropp, et al., Matrix concentration inequalities via the method of exchangeable pairs , The Annals of Probability 42
2014
Later among the works it cites.
2015
Closest in time.
Murat A Erdogdu and Andrea Montanari, Convergence rates of sub-sampled Newton methods , Advances in Neural Information Processing Systems 29-(NIPS-15), 2015
2015
Closest in time.
Murat A Erdogdu, Newton-Stein Method: A second order method for GLMs via Stein’s lemma , Advances in Neural Information Processing Systems 29-(NIPS-15), 2015
2015
Closest in time.