Fetching the paper…
Reading the bibliography…
First order methods, which solely rely on gradient information, are commonly used in diverse machine learning (ML) and data analysis (DA) applications.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
The elements of statistical learning
Jerome Friedman, Trevor Hastie, and Robert Tibshirani · 2001
Earlier work this paper cites.
Large scale online learning
Léon Bottou and Yann LeCun · 2004
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Scalable learning for object detection with gpu hardware
Adam Coates, Paul Baumstarck, Quoc Le, and Andrew Y Ng · 2009
Earlier work this paper cites.
Large-scale deep unsupervised learning using graphics processors
Rajat Raina, Anand Madhavan, and Andrew Y Ng · 2009
Earlier work this paper cites.
Blendenpik: Supercharging LAPACK’s least-squares solver
Haim Avron, Petar Maymounkov, and Sivan Toledo · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Randomized algorithms for matrices and data
Michael W Mahoney · 2011
Earlier work this paper cites.
On optimization methods for deep learning
Jiquan Ngiam, Adam Coates, Ahbik Lahiri, Bobby Prochnow, Quoc V Le, and Andrew Y Ng · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
Richard H. Byrd, Gillian M. Chin, Jorge Nocedal, and Yuchen Wu · 2012
Earlier work this paper cites.
Adaptive and stochastic algorithms for EIT and DC resistivity problems with piecewise constant solutions and many measurements
Kees van den Doel and Uri Ascher · 2012
Earlier work this paper cites.
Optimization for machine learning
Suvrit Sra, Sebastian Nowozin, and Stephen J Wright · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Deep learning with cots hpc systems
Adam Coates, Brody Huval, Tao Wang, David Wu, Bryan Catanzaro, and Ng Andrew · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Later among the works it cites.
Revisiting the Nyström method for improved large-scale machine learning
Alex Gittens and Michael W Mahoney · 2016
Later among the works it cites.
Sub-sampled Newton methods I: globally convergent algorithms
Farbod Roosta-Khorasani and Michael W Mahoney · 2016
Later among the works it cites.
Sub-sampled Newton methods II: Local convergence rates
Farbod Roosta-Khorasani and Michael W Mahoney · 2016
Later among the works it cites.
Sub-sampled newton methods with non-uniform sampling
Peng Xu, Jiyan Yang, Farbod Roosta-Khorasani, Christopher Ré, and Michael W Mahoney · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sudhir B Kylasa, Hasan Metin Aktulga, and Ananth Y Grama · 2014
Cited alongside, same era.
LSRN: A parallel iterative solver for strongly over-or underdetermined systems
Xiangrui Meng, Michael A Saunders, and Michael W Mahoney · 2014
Cited alongside, same era.
Data completion and stochastic algorithms for PDE inversion problems with many measurements
Farbod Roosta-Khorasani, Kees van den Doel, and Uri Ascher · 2014
Cited alongside, same era.
Stochastic algorithms for inverse problems involving PDEs and many measurements
Farbod Roosta-Khorasani, Kees van den Doel, and Uri Ascher · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Convergence rates of sub-sampled newton methods
Murat A. Erdogdu and Andrea Montanari · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Implementing randomized matrix algorithms in parallel and distributed environments
Jiyan Yang, Xiangrui Meng, and Michael W Mahoney · 2016
Later among the works it cites.
An Investigation of Newton-Sketch and Subsampled Newton Methods
Albert S Berahas, Raghu Bollapragada, and Jorge Nocedal · 2017
Later among the works it cites.
Newton-Type Methods for Non-Convex Optimization Under Inexact Hessian Information
Peng Xu, Farbod Roosta-Khorasani, and Michael W. Mahoney · 2017
Later among the works it cites.
Second-Order Optimization for Non-Convex Machine Learning: An Empirical Study
Peng Xu, Farbod Roosta-Khorasani, and Michael W. Mahoney · 2017
Later among the works it cites.
Newton-cg cuda implementation download (scripts/code/tensorflow-python-scripts)
Sudhir B Kylasa · 2018
Closest in time.
Gpu accelerated sub-sampled newton methods
Sudhir B Kylasa, Farbod Roosta-Khorasani, Michael W. Mahoney, and Ananth Y Grama · 2018
Closest in time.
Uci machine learning repository
UCI · 2018
Closest in time.