Fetching the paper…
Reading the bibliography…
Stochastic optimization algorithms with variance reduction have proven successful for minimizing large finite sums of functions.
Convex analysis and minimization algorithms I: Fundamentals
J.-B. Hiriart-Urruty and C. Lemaréchal · 1993
Earlier work this paper cites.
Transformation Invariance in Pattern Recognition — Tangent Distance and Tangent Propagation
P. Y. Simard, Y. A. LeCun, J. S. Denker, and B. Victorri · 1998
Earlier work this paper cites.
A Gene-Expression Signature as a Predictor of Survival in Breast Cancer
M. J. van de Vijver et al · 2002
Earlier work this paper cites.
Introductory Lectures on Convex Optimization
Y. Nesterov · 2004
Earlier work this paper cites.
Training invariant support vector machines using selective sampling
G. Loosli, S. Canu, and L. Bottou · 2007
Earlier work this paper cites.
Efficient online and batch learning using forward backward splitting
J. C. Duchi and Y. Singer · 2009
Earlier work this paper cites.
Robust Stochastic Approximation Approach to Stochastic Programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Stability selection
N. Meinshausen and P. Bühlmann · 2010
Earlier work this paper cites.
Dual averaging methods for regularized stochastic learning and online optimization
L. Xiao · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. Bach and E. Moulines · 2011
Earlier work this paper cites.
An Analysis of Single-Layer Networks in Unsupervised Feature Learning
A. Coates, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
Privacy aware learning
J. C. Duchi, M. I. Jordan, and M. J. Wainwright · 2012
Cited alongside, same era.
S. Lacoste-Julien, M. Schmidt, and F. Bach · 2012
Cited alongside, same era.
Invariant scattering convolution networks
J. Bruna and S. Mallat · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2013
Cited alongside, same era.
Learning with marginalized corrupted features
L. van der Maaten, M. Chen, S. Tyree, and K. Q. Weinberger · 2013
Variance Reduced Stochastic Gradient Descent with Neighbors
T. Hofmann, A. Lucchi, S. Lacoste-Julien, and B. McWilliams · 2015
Later among the works it cites.
A Universal Catalyst for First-Order Optimization
H. Lin, J. Mairal, and Z. Harchaoui · 2015
Later among the works it cites.
Incremental Majorization-Minimization Optimization with Application to Large-Scale Machine Learning
J. Mairal · 2015
Later among the works it cites.
Exploiting the Structure: Stochastic Gradient Methods Using Raw Clusters
Z. Allen-Zhu, Y. Yuan, and K. Sridharan · 2016
Closest in time.
Optimization Methods for Large-Scale Machine Learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Closest in time.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
Finito: A faster, permutable incremental gradient method for big data problems
A. Defazio, J. Domke, and T. S. Caetano · 2014
Cited alongside, same era.
Transformation pursuit for image classification
M. Paulin, J. Revaud, Z. Harchaoui, F. Perronnin, and C. Schmid · 2014
Cited alongside, same era.
Altitude Training: Strong Bounds for Single-layer Dropout
S. Wager, W. Fithian, S. Wang, and P. Liang · 2014
Cited alongside, same era.
A proximal stochastic gradient method with progressive variance reduction
L. Xiao and T. Zhang · 2014
Cited alongside, same era.
SGD with Variance Reduction beyond Empirical Risk Minimization
M. Achab, A. Guilloux, S. Gaïffas, and E. Bacry · 2015
Cited alongside, same era.
Closest in time.
End-to-End Kernel Learning with Supervised Convolutional Kernel Networks
J. Mairal · 2016
Closest in time.
SDCA without Duality, Regularization, and Individual Convexity
S. Shalev-Shwartz · 2016
Closest in time.
Improving the robustness of deep neural networks via stability training
S. Zheng, Y. Song, T. Leung, and I. Goodfellow · 2016
Closest in time.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2017
Closest in time.
An optimal randomized incremental gradient method
G. Lan and Y. Zhou · 2017
Closest in time.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Closest in time.