Fetching the paper…
Reading the bibliography…
Lecture notes on optimization for machine learning, derived from a course at Princeton University and tutorials given in MLSS, Buenos Aires, as well as Simons Foundation, Berkeley.
Computing machinery and intelligence
A. M. Turing · 1950
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
An algorithm for quadratic programming
M. Frank and P. Wolfe · 1956
Earlier work this paper cites.
Approximation to bayes risk in repeated play
James Hannan · 1957
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Arkadi S. Nemirovski and David B. Yudin · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Interior Point Polynomial Algorithms in Convex Programming
Y. E. Nesterov and A. S. Nemirovskii · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
Jyrki Kivinen and Manfred K. Warmuth · 1997
Earlier work this paper cites.
Convex Analysis
R.T. Rockafellar · 1997
Earlier work this paper cites.
General convergence results for linear discriminant updates
A .J. Grove, N. Littlestone, and D. Schuurmans · 2001
Earlier work this paper cites.
Relative loss bounds for multidimensional regression problems
Jyrki Kivinen and Manfred K. Warmuth · 2001
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Interior point polynomial time methods in convex programming, 2004
A.S. Nemirovskii · 2004
Earlier work this paper cites.
Interior point polynomial time methods in convex programming
AS Nemirovskii · 2004
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Y. Nesterov · 2004
Earlier work this paper cites.
Learning with Matrix Factorizations
Nathan Srebro · 2004
Earlier work this paper cites.
Efficient algorithms for online decision problems
Adam Kalai and Santosh Vempala · 2005
Earlier work this paper cites.
Fast maximum margin matrix factorization for collaborative prediction
Jasson D. M. Rennie and Nathan Srebro · 2005
Earlier work this paper cites.
Convex Analysis and Nonlinear Optimization: Theory and Examples
J.M. Borwein and A.S. Lewis · 2006
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolò Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
Online Learning: Theory, Algorithms, and Applications
Shai Shalev-Shwartz · 2007
Earlier work this paper cites.
A primal-dual perspective of online learning algorithms
Shai Shalev-Shwartz and Yoram Singer · 2007
Earlier work this paper cites.
Competing in the dark: An efficient algorithm for bandit linear optimization
Jacob Abernethy, Elad Hazan, and Alexander Rakhlin · 2008
Earlier work this paper cites.
Extracting certainty from uncertainty: Regret bounded by variation in costs
Elad Hazan and Satyen Kale · 2008
Earlier work this paper cites.
Exact matrix completion via convex optimization
E. Candes and B. Recht · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2010
Earlier work this paper cites.
A simple algorithm for nuclear norm regularized problems
Martin Jaggi and Marek Sulovský · 2010
Cited alongside, same era.
Practical large-scale optimization for max-norm regularization
J. Lee, B. Recht, R. Salakhutdinov, N. Srebro, and J. A. Tropp · 2010
Cited alongside, same era.
Adaptive bound optimization for online convex optimization
H. Brendan McMahan and Matthew J. Streeter · 2010
Cited alongside, same era.
New adaptive algorithms for online classification
Francesco Orabona and Koby Crammer · 2010
Cited alongside, same era.
Compressive sensing and structured random matrices
Holger Rauhut · 2010
Cited alongside, same era.
Collaborative filtering in a non-uniform world: Learning with the weighted trace norm
R. Salakhutdinov and N. Srebro · 2010
Cited alongside, same era.
Linear convergence with condition number independent access of full gradients
Lijun Zhang, Mehrdad Mahdavi, and Rong Jin · 2013
Later among the works it cites.
Linear coupling: An ultimate unification of gradient and mirror descent
Zeyuan Allen-Zhu and Lorenzo Orecchia · 2014
Later among the works it cites.
Distributed frank-wolfe algorithm: A unified framework for communication-efficient sparse learning
Aurélien Bellet, Yingyu Liang, Alireza Bagheri Garakani, Maria-Florina Balcan, and Fei Sha · 2014
Later among the works it cites.
Bayesian optimization with inequality constraints
Jacob R. Gardner, Matt J. Kusner, Zhixiang Eddie Xu, Kilian Q. Weinberger, and John P. Cunningham · 2014
Later among the works it cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Algorithms for hyper-parameter optimization
James S. Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Approximating semidefinite programs in sublinear time
Dan Garber and Elad Hazan · 2011
Cited alongside, same era.
Beyond the regret minimization barrier: an optimal algorithm for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2011
Cited alongside, same era.
Large-scale convex minimization with a low-rank constraint
Shai Shalev-Shwartz, Alon Gonen, and Ohad Shamir · 2011
Cited alongside, same era.
Pegasos: primal estimated sub-gradient solver for svm
Shai Shalev-Shwartz, Yoram Singer, Nathan Srebro, and Andrew Cotter · 2011
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Later among the works it cites.
Analysis of Boolean Functions
Ryan O’Donnell · 2014
Later among the works it cites.
Input warping for bayesian optimization of non-stationary functions
Jasper Snoek, Kevin Swersky, Richard S. Zemel, and Ryan P. Adams · 2014
Later among the works it cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Introduction to online convex optimization
Elad Hazan · 2016
Later among the works it cites.
Non-stochastic best arm identification and hyperparameter optimization
Kevin G. Jamieson and Ameet Talwalkar · 2016
Later among the works it cites.
Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization
L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar · 2016
Later among the works it cites.
Embracing the random
Benjamin Recht · 2016
Later among the works it cites.
The news on auto-tuning
Benjamin Recht · 2016
Later among the works it cites.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Later among the works it cites.
Second-order stochastic optimization for machine learning in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2017
Later among the works it cites.
A unified approach to adaptive regularization in online and stochastic optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2017
Later among the works it cites.
Efficient hyperparameter optimization for deep learning algorithms using deterministic RBF surrogates
Ilija Ilievski, Taimoor Akhtar, Jiashi Feng, and Christine Annette Shoemaker · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Later among the works it cites.
The case for full-matrix adaptive regularization
Naman Agarwal, Brian Bullins, Xinyi Chen, Elad Hazan, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Later among the works it cites.
Optimal adaptive and accelerated stochastic gradient descent
Qi Deng, Yi Cheng, and Guanghui Lan · 2018
Later among the works it cites.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Later among the works it cites.
Hyperparameter optimization: A spectral approach
Elad Hazan, Adam Klivans, and Yang Yuan · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2018
Later among the works it cites.
Memory-efficient adaptive optimization for large-scale learning
Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer · 2019
Closest in time.
Extreme tensoring for low-memory preconditioning
Xinyi Chen, Naman Agarwal, Elad Hazan, Cyril Zhang, and Yi Zhang · 2019
Closest in time.
Revisiting the polyak step size
Elad Hazan and Sham Kakade · 2019
Closest in time.