Fetching the paper…
Reading the bibliography…
Incremental gradient (IG) methods, such as stochastic gradient descent and its variants are commonly used for large scale optimization in machine learning.
On a stochastic approximation method
Chung, K. L. et al · 1954
Earlier work this paper cites.
Analysis of an approximate gradient projection method with applications to the backpropagation algorithm
Zhi-Quan, L. and Paul, T · 1954
Earlier work this paper cites.
Accelerated greedy algorithms for maximizing submodular set functions
Minoux, M · 1978
Earlier work this paper cites.
An analysis of approximations for maximizing submodular set functions—i
Nemhauser, G., Wolsey, L., and Fisher, M · 1978
Earlier work this paper cites.
An analysis of the greedy algorithm for the submodular set covering problem
Wolsey, L. A · 1982
Earlier work this paper cites.
Clustering by means of medoids in statistical data analysis based on the, 1987
Kaufman, L., Rousseeuw, P., and Dodge, Y · 1987
Earlier work this paper cites.
Serial and parallel backpropagation convergence via nonmonotone perturbed minimization
Mangasariany, O. and Solodovy, M · 1994
Earlier work this paper cites.
Incremental least squares methods and the extended kalman filter
Bertsekas, D. P · 1996
Earlier work this paper cites.
Incremental gradient algorithms with stepsizes bounded away from zero
Solodov, M. V · 1998
Earlier work this paper cites.
An incremental gradient (-projection) method with momentum term and adaptive stepsize rule
Tseng, P · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N · 1999
Earlier work this paper cites.
Convergence rate of incremental subgradient algorithms
Nedić, A. and Bertsekas, D · 2001
Earlier work this paper cites.
Approximating extent measures of points
Agarwal, P. K., Har-Peled, S., and Varadarajan, K. R · 2004
Earlier work this paper cites.
On coresets for k-means and k-median clustering
Har-Peled, S. and Mazumdar, S · 2004
Earlier work this paper cites.
Graph-based submodular selection for extractive summarization
Lin, H., Bilmes, J., and Xie, S · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
Learning mixtures of submodular shells with application to document summarization
Lin, H. and Bilmes, J. A · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Roux, N. L., Schmidt, M., and Bach, F. R · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Zeiler, M. D · 2012
Cited alongside, same era.
A universal catalyst for first-order optimization
Lin, H., Mairal, J., and Harchaoui, Z · 2015
Later among the works it cites.
Submodularity in data subset selection and active learning
Wei, K., Iyer, R., and Bilmes, J · 2015
Later among the works it cites.
Exploiting the structure: Stochastic gradient methods using raw clusters
Allen-Zhu, Z., Yuan, Y., and Sridharan, K · 2016
Later among the works it cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Later among the works it cites.
Fast distributed submodular cover: Public-private data summarization
Mirzasoleiman, B., Zadimoghaddam, M., and Karbasi, A · 2016
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Allen-Zhu, Z · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Cited alongside, same era.
Iterative row sampling
Li, M., Miller, G. L., and Peng, R · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shalev-Shwartz, S. and Zhang, T · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
A proximal stochastic gradient method with progressive variance reduction
Xiao, L. and Zhang, T · 2014
Cited alongside, same era.
Un-regularizing: approximate proximal point and faster stochastic algorithms for empirical risk minimization
Frostig, R., Ge, R., Kakade, S., and Sidford, A · 2015
Cited alongside, same era.
Later among the works it cites.
Input sparsity time low-rank approximation via ridge leverage score sampling
Cohen, M. B., Musco, C., and Musco, C · 2017
Later among the works it cites.
On the convergence rate of incremental aggregated gradient algorithms
Gurbuzbalaban, M., Ozdaglar, A., and Parrilo, P. A · 2017
Later among the works it cites.
Training gaussian mixture models at scale via coresets
Lucic, M., Faulkner, M., Krause, A., and Feldman, D · 2017
Later among the works it cites.
Recursive sampling for the nystrom method
Musco, C. and Musco, C · 2017
Later among the works it cites.
Bayesian coreset construction via greedy iterative geodesic ascent
Campbell, T. and Broderick, T · 2018
Later among the works it cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Later among the works it cites.
The importance of better models in stochastic optimization
Asi, H. and Duchi, J. C · 2019
Closest in time.
On the ineffectiveness of variance reduced optimization for deep learning
Defazio, A. and Bottou, L · 2019
Closest in time.
Energy and policy considerations for deep learning in nlp
Strubell, E., Ganesh, A., and McCallum, A · 2019
Closest in time.