Fetching the paper…
Reading the bibliography…
First-order optimization methods, such as stochastic gradient descent (SGD) and its variants, are widely used in machine learning applications due to their simplicity and low per-iteration costs.
Invex functions and constrained local minima
BD Craven · 1981
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
The elements of statistical learning
Jerome Friedman, Trevor Hastie, and Robert Tibshirani · 2001
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, Jonathan Eckstein, et al · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Kevin P Murphy · 2012
Earlier work this paper cites.
Optimization for machine learning
Suvrit Sra, Sebastian Nowozin, and Stephen J Wright · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Searching for exotic particles in high-energy physics with deep learning
Pierre Baldi, Peter Sadowski, and Daniel Whiteson · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Disco: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Xiao Lin · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
Revisiting distributed synchronous sgd
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Cited alongside, same era.
DynaNewton-Accelerating Newton’s Method for Machine Learning
Hadi Daneshmand, Aurelien Lucchi, and Thomas Hofmann · 2016
Cited alongside, same era.
Accurate, large minibatch sgd: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Later among the works it cites.
Second-Order Optimization for Non-Convex Machine Learning: An Empirical Study
Peng Xu, Farbod Roosta-Khorasani, and Michael W. Mahoney · 2017
Later among the works it cites.
Adaptive relaxed admm: Convergence theory and practical implementation
Zheng Xu, Mário AT Figueiredo, Xiaoming Yuan, Christoph Studer, and Tom Goldstein · 2017
Later among the works it cites.
Adaptive consensus admm for distributed optimization
Zheng Xu, Gavin Taylor, Hao Li, Mario Figueiredo, Xiaoming Yuan, and Tom Goldstein · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter H Jin, Qiaochu Yuan, Forrest Iandola, and Kurt Keutzer · 2016
Cited alongside, same era.
AIDE: Fast and communication efficient distributed optimization
Sashank J Reddi, Jakub Konečnỳ, Peter Richtárik, Barnabás Póczós, and Alex Smola · 2016
Cited alongside, same era.
First-Order Methods in Optimization
Amir Beck · 2017
Cited alongside, same era.
An Investigation of Newton-Sketch and Subsampled Newton Methods
Albert S Berahas, Raghu Bollapragada, and Jorge Nocedal · 2017
Cited alongside, same era.
Celestine Dünner, Aurelien Lucchi, Matilde Gargiani, An Bian, Thomas Hofmann, and Martin Jaggi · 2018
Closest in time.
GIANT: Globally Improved Approximate Newton Method for Distributed Optimization
Shusen Wang, Farbod Roosta-Khorasani, Peng Xu, and Michael W Mahoney · 2018
Closest in time.
DINGO: Distributed Newton-Type Method for Gradient-Norm Optimization
Rixon Crane and Fred Roosta · 2019
Closest in time.
GPU Accelerated Sub-Sampled Newton’s Method for Convex Classification Problems
Sudhir Kylasa, Farbod Roosta-Khorasani, Michael W Mahoney, and Ananth Grama · 2019
Closest in time.
Sub-sampled Newton methods
Farbod Roosta-Khorasani and Michael W Mahoney · 2019
Closest in time.