Fetching the paper…
Reading the bibliography…
This paper studies the empirical efficacy and benefits of using projection-free first-order methods in the form of Conditional Gradients, a.k.a.
An algorithm for quadratic programming
M. Frank, P. Wolfe, et al · 1956
Earlier work this paper cites.
Constrained minimization methods
E. S. Levitin and B. T. Polyak · 1966
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o (1/k2̂)
Y. Nesterov · 1983
Earlier work this paper cites.
Comparing biases for minimal network construction with back-propagation
S. J. Hanson and L. Y. Pratt · 1989
Earlier work this paper cites.
Linear best approximation using a class of polyhedral norms
G. A. Watson · 1992
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Efficient projections onto the l 1 l_{1} -ball for learning in high dimensions
J. Duchi, S. Shalev-Shwartz, Y. Singer, and T. Chandra · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Mnist handwritten digit database
Y. LeCun, C. Cortes, and C. Burges · 2010
Earlier work this paper cites.
Y. Chen and X. Ye · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
Online linear optimization over permutations
S. Yasutake, K. Hatano, S. Kijima, E. Takimoto, and M. Takeda · 2011
Earlier work this paper cites.
Revisiting frank-wolfe: Projection-free sparse convex optimization
M. Jaggi · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
W. Wang and M. A. Carreira-Perpiñán · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Projection-free online optimization with stochastic gradient: From convexity to submodularity
L. Chen, C. Harshaw, H. Hassani, and A. Karbasi · 2018
Later among the works it cites.
SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
C. Fang, C. J. Li, Z. Lin, and T. Zhang · 2018
Later among the works it cites.
Conditional gradient method for stochastic submodular maximization: Closing the gap
A. Mokhtari, H. Hassani, and A. Karbasi · 2018
Later among the works it cites.
Momentum-based variance reduction in non-convex SGD
A. Cutkosky and F. Orabona · 2019
Later among the works it cites.
On the ineffectiveness of variance reduced optimization for deep learning
A. Defazio and L. Bottou · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Variance-reduced and projection-free stochastic optimization
E. Hazan and H. Luo · 2016
Cited alongside, same era.
Conditional gradient sliding for convex optimization
G. Lan and Y. Zhou · 2016
Cited alongside, same era.
Efficient bregman projections onto the permutahedron and related polytopes
C. H. Lim and S. J. Wright · 2016
Cited alongside, same era.
Stochastic frank-wolfe methods for nonconvex optimization
S. J. Reddi, S. Sra, B. Póczos, and A. Smola · 2016
Cited alongside, same era.
Linear convergence of stochastic frank wolfe variants
D. Goldfarb, G. Iyengar, and C. Zhou · 2017
Cited alongside, same era.
Conditional accelerated lazy stochastic gradient descent
G. Lan, S. Pokutta, Y. Zhou, and D. Zink · 2017
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Cited alongside, same era.
U. Evci, T. Gale, J. Menick, P. S. Castro, and E. Elsen · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Later among the works it cites.
Complexities in projection-free stochastic non-convex minimization
Z. Shen, C. Fang, P. Zhao, J. Huang, and H. Qian · 2019
Later among the works it cites.
Stochastic recursive gradient-based methods for projection-free online learning
J. Xie, Z. Shen, C. Zhang, H. Qian, and B. Wang · 2019
Later among the works it cites.
Conditional gradient methods via stochastic path-integrated differential estimator
A. Yurtsever, S. Sra, and V. Cevher · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
L. Biewald · 2020
Closest in time.
Projection-free adaptive gradients for large-scale optimization
C. W. Combettes, C. Spiegel, and S. Pokutta · 2020
Closest in time.
Stochastic conditional gradient methods: From convex minimization to submodular maximization
A. Mokhtari, H. Hassani, and A. Karbasi · 2020
Closest in time.
Stochastic frank-wolfe for constrained finite-sum minimization
G. Négiar, G. Dresdner, A. Tsai, L. E. Ghaoui, F. Locatello, and F. Pedregosa · 2020
Closest in time.
Descending through a crowded valley–benchmarking deep learning optimizers
R. M. Schmidt, F. Schneider, and P. Hennig · 2020
Closest in time.
One sample stochastic frank-wolfe
M. Zhang, Z. Shen, A. Mokhtari, H. Hassani, and A. Karbasi · 2020
Closest in time.