Fetching the paper…
Reading the bibliography…
We present an approach towards convex optimization that relies on a novel scheme which converts online adaptive algorithms into offline methods.
Problem complexity and method efficiency in optimization
Nemirovskii, Arkadii, Yudin, David Borisovich, and Dawson, ER · 1983
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o (1/k2)
Nesterov, Yurii · 1983
Earlier work this paper cites.
Minimization methods for nonsmooth convex and quasiconvex functions
Nesterov, Yu E · 1984
Earlier work this paper cites.
Numerical optimization
Wright, Stephen and Nocedal, Jorge · 1999
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Cesa-Bianchi, Nicolo, Conconi, Alex, and Gentile, Claudio · 2004
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Hazan, Elad, Agarwal, Amit, and Kale, Satyen · 2007
Earlier work this paper cites.
Large deviations of vector-valued martingales in 2-smooth normed spaces
Juditsky, Anatoli B and Nemirovski, Arkadi S · 2008
Earlier work this paper cites.
Markov chains and mixing times
Levin, David Asher, Peres, Yuval, and Wilmer, Elizabeth Lee · 2009
Earlier work this paper cites.
Lecture notes in multivariate analysis, dimensionality reduction, and spectral methods
Kakade, Sham · 2010
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
McMahan, H Brendan and Streeter, Matthew · 2010
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Cotter, Andrew, Shamir, Ohad, Srebro, Nati, and Sridharan, Karthik · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Cited alongside, same era.
Sublinear optimization for machine learning
Clarkson, Kenneth L, Hazan, Elad, and Woodruff, David P · 2012
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
Dekel, Ofer, Gilad-Bachrach, Ran, Shamir, Ohad, and Xiao, Lin · 2012
Cited alongside, same era.
Linear regression with limited observation
Hazan, Elad and Koren, Tomer · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
Accelerated mini-batch stochastic dual coordinate ascent
Shalev-Shwartz, Shai and Zhang, Tong · 2013
Later among the works it cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Later among the works it cites.
Efficient mini-batch training for stochastic optimization
Li, Mu, Zhang, Tong, Chen, Yuqiang, and Smola, Alexander J · 2014
Later among the works it cites.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, Elad, Levy, Kfir, and Shalev-Shwartz, Shai · 2015
Later among the works it cites.
A universal catalyst for first-order optimization
Lin, Hongzhou, Mairal, Julien, and Harchaoui, Zaid · 2015
Later among the works it cites.
Universal gradient methods for convex optimization problems
Nesterov, Yu · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adadelta: an adaptive learning rate method
Zeiler, Matthew D · 2012
Cited alongside, same era.
Means and their Inequalities , volume 31
Bullen, Peter S, Mitrinovic, Dragoslav S, and Vasic, M · 2013
Cited alongside, same era.
Gradient methods for minimizing composite functions
Nesterov, Yu · 2013
Cited alongside, same era.
Later among the works it cites.
Takáč, Martin, Richtárik, Peter, and Srebro, Nathan · 2015
Later among the works it cites.
Parallelizing stochastic approximation through mini-batching and tail-averaging
Jain, Prateek, Kakade, Sham M, Kidambi, Rahul, Netrapalli, Praneeth, and Sidford, Aaron · 2016
Later among the works it cites.
The power of normalization: Faster evasion of saddle points
Levy, Kfir Y · 2016
Later among the works it cites.