Fetching the paper…
Reading the bibliography…
We develop and analyze a procedure for gradient-based optimization that we refer to as stochastically controlled stochastic gradient (SCSG).
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Ralph Bradley and Milton Terry · 1952
Earlier work this paper cites.
Robust Statistics
Peter J Huber · 1981
Earlier work this paper cites.
Generalized Linear Models
Peter McCullagh and John A Nelder · 1989
Earlier work this paper cites.
Dao Li Zhu and Patrice Marcotte · 1996
Earlier work this paper cites.
An improved approximation algorithm for multiway cut
Gruia Călinescu, Howard Karloff, and Yuval Rabani · 1998
Earlier work this paper cites.
Asymptotic Statistics
Aad W Van der Vaart · 1998
Earlier work this paper cites.
Distance metric learning with application to clustering with side-information
Eric P Xing, Andrew Y Ng, Michael I Jordan, and Stuart Russell · 2002
Earlier work this paper cites.
MM algorithms for generalized Bradley-Terry models
David R Hunter · 2004
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2004
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
Kilian Q Weinberger, John Blitzer, and Lawrence Saul · 2006
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Ohad Shamir · 2011
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Le Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux · 2012
Cited alongside, same era.
Proximal stochastic dual coordinate ascent
Shai Shalev-Shwartz and Tong Zhang · 2012
Cited alongside, same era.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, Martin J Wainwright, and John C Duchi · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
A lower bound for the optimization of finite sums
Alekh Agarwal and Leon Bottou · 2014
Cited alongside, same era.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Later among the works it cites.
Competing with the empirical risk minimizer in a single pass
Roy Frostig, Rong Ge, Sham M Kakade, and Aaron Sidford · 2015
Later among the works it cites.
Stop wasting my gradients: Practical SVRG
Reza Harikandeh, Mohamed Osama Ahmed, Alim Virani, Mark Schmidt, Jakub Konečnỳ, and Scott Sallinen · 2015
Later among the works it cites.
Variance reduced stochastic gradient descent with neighbors
Thomas Hofmann, Aurelien Lucchi, Simon Lacoste-Julien, and Brian McWilliams · 2015
Later among the works it cites.
Federated optimization: Distributed optimization beyond the datacenter
Jakub Konečnỳ, Brendan McMahan, and Daniel Ramage · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeyuan Allen-Zhu and Lorenzo Orecchia · 2014
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takác, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I Jordan · 2014
Cited alongside, same era.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Cited alongside, same era.
An accelerated proximal coordinate gradient method
Qihang Lin, Zhaosong Lu, and Lin Xiao · 2014
Cited alongside, same era.
Stochastic proximal gradient descent with acceleration techniques
Atsushi Nitanda · 2014
Cited alongside, same era.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2014
Cited alongside, same era.
Jason Lee, Tengyu Ma, and Qihang Lin · 2015
Later among the works it cites.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Later among the works it cites.
Accelerated stochastic gradient descent for minimizing finite sums
Atsushi Nitanda · 2015
Later among the works it cites.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank J. Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczos, and Alex J. Smola · 2015
Later among the works it cites.
Stochastic primal-dual coordinate method for regularized empirical risk minimization
Yuchen Zhang and Lin Xiao · 2015
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2016
Closest in time.
Starting small–learning with adaptive sample sizes
Hadi Daneshmand, Aurelien Lucchi, and Thomas Hofmann · 2016
Closest in time.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Closest in time.
Tight complexity bounds for optimizing composite objectives
Blake Woodworth and Nathan Srebro · 2016
Closest in time.