Fetching the paper…
Reading the bibliography…
The complexity in large-scale optimization can lie in both handling the objective function and handling the constraint set.
An algorithm for quadratic programming
M. Frank and P. Wolfe · 1956
Earlier work this paper cites.
Constrained minimization methods
E. S. Levitin and B. T. Polyak · 1966
Earlier work this paper cites.
Introduction to Optimization
B. T. Polyak · 1987
Earlier work this paper cites.
Text categorization based on regularized linear classification methods
T. Zhang and F. J. Oles · 2001
Earlier work this paper cites.
RCV1: A new benchmark collection for text categorization research
D. D. Lewis, Y. Yang, T. G. Rose, and F. Li · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. B. McMahan and M. Streeter · 2010
Earlier work this paper cites.
The Million Song dataset
T. Bertin-Mahieux, D. P. W. Ellis, B. Whitman, and P. Lamere · 2011
Earlier work this paper cites.
LIBSVM: A library for support vector machines
C.-C. Chang and C.-J. Lin · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. C. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng · 2012
Earlier work this paper cites.
Lecture 6e – rmsprop: divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
AdaDelta: An adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
Revisiting Frank-Wolfe: Projection-free sparse convex optimization
M. Jaggi · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Linear convergence with condition number independent access of full gradients
L. Zhang, M. Mahdavi, and R. Jin · 2013
Cited alongside, same era.
Hybrid conditional gradient-smoothing algorithms with applications to sparse and low rank regularization
A. Argyriou, M. Signoretto, and J. A. K. Suykens · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Introductory lectures on stochastic optimization
J. C. Duchi · 2018
Later among the works it cites.
SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
C. Fang, C. J. Li, Z. Lin, and T. Zhang · 2018
Later among the works it cites.
On the convergence of Adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Later among the works it cites.
Momentum-based variance reduction in non-convex SGD
A. Cutkosky and F. Orabona · 2019
Later among the works it cites.
On the ineffectiveness of variance reduced optimization for deep learning
A. Defazio and L. Bottou · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, Y. Liu, and X. Sun · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Variance-reduced and projection-free stochastic optimization
E. Hazan and H. Luo · 2016
Cited alongside, same era.
Convergence rate of Frank-Wolfe for non-convex objectives
S. Lacoste-Julien · 2016
Cited alongside, same era.
Conditional gradient sliding for convex optimization
G. Lan and Y. Zhou · 2016
Cited alongside, same era.
Stochastic Frank-Wolfe methods for nonconvex optimization
S. J. Reddi, S. Sra, B. Póczos, and A. Smola · 2016
Cited alongside, same era.
Improving generalization performance by switching from Adam to SGD
N. S. Keskar and R. Socher · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Cited alongside, same era.
Complexities in projection-free stochastic non-convex minimization
Z. Shen, C. Fang, P. Zhao, J. Huang, and H. Qian · 2019
Later among the works it cites.
Conditional gradient methods via stochastic path-integrated differential estimator
A. Yurtsever, S. Sra, and V. Cevher · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
L. Biewald · 2020
Closest in time.
Boosting Frank-Wolfe by chasing gradients
C. W. Combettes and S. Pokutta · 2020
Closest in time.
Stochastic Frank-Wolfe for constrained finite-sum minimization
G. Négiar, G. Dresdner, A. Y.-T. Tsai, L. El Ghaoui, F. Locatello, R. M. Freund, and F. Pedregosa · 2020
Closest in time.
Efficient projection-free online methods with stochastic recursive gradient
J. Xie, Z. Shen, C. Zhang, H. Qian, and B. Wang · 2020
Closest in time.
Complexity of linear minimization and projection on some sets
C. W. Combettes and S. Pokutta · 2021
Closest in time.