Fetching the paper…
Reading the bibliography…
Stochastic gradient methods for machine learning and optimization problems are usually analyzed assuming data points are sampled \emph{with} replacement.
Statistical learning theory
V. Vapnik · 1998
Earlier work this paper cites.
Nonlinear programming
D. Bertsekas · 1999
Earlier work this paper cites.
Convergence rate of incremental subgradient algorithms
A. Nedić and D. Bertsekas · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. Bartlett and S. Mendelson · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
L. Bottou · 2009
Earlier work this paper cites.
Transductive rademacher complexity and its applications
R. El-Yaniv and D. Pechyony · 2009
Earlier work this paper cites.
Note on sampling without replacing from a finite collection of matrices
D. Gross and V. Nesme · 2010
Earlier work this paper cites.
Dual averaging methods for regularized stochastic learning and online optimization
L. Xiao · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, A. Smola, and L. Li · 2010
Earlier work this paper cites.
A reliable effective terascale linear learning system
A. Agarwal, O. Chapelle, M. Dudík, and J. Langford · 2011
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
R. Bekkerman, M. Bilenko, and J. Langford · 2011
Earlier work this paper cites.
Distributed optimization and statistical learning via ADMM
S.P. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2011
Earlier work this paper cites.
Online learning and online convex optimization
S. Shalev-Shwartz · 2011
Cited alongside, same era.
Distributed learning, communication complexity and privacy
M.-F. Balcan, A. Blum, S. Fine, and Y. Mansour · 2012
Cited alongside, same era.
Stochastic gradient descent tricks
L. Bottou · 2012
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Cited alongside, same era.
Dual averaging for distributed optimization: Convergence analysis and network scaling
J. Duchi, A. Agarwal, and M. Wainwright · 2012
Cited alongside, same era.
Towards a unified architecture for in-rdbms analytics
X. Feng, A. Kumar, B. Recht, and C. Ré · 2012
Cited alongside, same era.
Trading computation for communication: Distributed SDCA
T. Yang · 2013
Later among the works it cites.
Communication-efficient algorithms for statistical optimization
Y. Zhang, J. Duchi, and M. Wainwright · 2013
Later among the works it cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Later among the works it cites.
Finito: A faster, permutable incremental gradient method for big data problems
A. J Defazio, T. Caetano, and J. Domke · 2014
Later among the works it cites.
Communication-efficient distributed dual coordinate ascent
M. Jaggi, V. Smith, M. Takác, J. Terhorst, S. Krishnan, T. Hofmann, and M. Jordan · 2014
Later among the works it cites.
Understanding machine learning: From Theory to Algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Lacoste-Julien, M. Schmidt, and F. Bach · 2012
Cited alongside, same era.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Cited alongside, same era.
Beneath the valley of the noncommutative arithmetic-geometric mean inequality: conjectures, case-studies, and consequences
B. Recht and C. Ré · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
N. Roux, M. Schmidt, and F. Bach · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Semi-stochastic gradient descent methods
J. Konečnỳ and P. Richtárik · 2013
Cited alongside, same era.
S. Shalev-Shwartz and S. Ben-David · 2014
Later among the works it cites.
On distributed stochastic optimization and learning
O. Shamir and N. Srebro · 2014
Later among the works it cites.
Communication-efficient distributed optimization using an approximate newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Later among the works it cites.
Convex optimization: Algorithms and complexity
S. Bubeck · 2015
Later among the works it cites.
Why random reshuffling beats stochastic gradient descent
M. Gürbüzbalaban, A. Ozdaglar, and P. Parrilo · 2015
Later among the works it cites.
Introduction to online convex optimization
E. Hazan · 2015
Later among the works it cites.
Distributed stochastic variance reduced gradient methods
J. Lee, T. Ma, and Q. Lin · 2015
Later among the works it cites.
An introduction to matrix concentration inequalities
J. Tropp · 2015
Later among the works it cites.
Communication-efficient distributed optimization of self-concordant empirical loss
Y. Zhang and L. Xiao · 2015
Later among the works it cites.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2016
Closest in time.