Training deep networks without learning rates through coin betting
Francesco Orabona and Tatiana Tommasi · 2017
Later among the works it cites.
Reflections on random kitchen sinks, 2017
Ali Rahimi and Ben Recht · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Later among the works it cites.
Stochastic gradient descent: Going as fast as possible but not faster
Original
Alice Schoenauer-Sebag, Marc Schoenauer, and Michèle Sebag · 2017
Later among the works it cites.
Stabilizing adversarial nets with prediction methods
Original
Abhay Yadav, Sohil Shah, Zheng Xu, David Jacobs, and Tom Goldstein · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Yellowfin and the art of momentum tuning
Original
Jian Zhang and Ioannis Mitliagkas · 2017
Later among the works it cites.
On exponential convergence of SGD in non-convex over-parametrized learning
Original
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Later among the works it cites.
A progressive batching l-bfgs method for machine learning
Raghu Bollapragada, Jorge Nocedal, Dheevatsa Mudigere, Hao-Jun Shi, and Ping Tak Peter Tang · 2018
Later among the works it cites.
On the linear convergence of the stochastic gradient method with constant step-size
Volkan Cevher and Bang Công Vũ · 2018
Later among the works it cites.
Accelerating stochastic gradient descent for least squares regression
Prateek Jain, Sham Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Later among the works it cites.
An alternative view: When does SGD escape local minima?
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Later among the works it cites.
Just interpolate: Kernel" ridgeless" regression can generalize
Original
Tengyuan Liang and Alexander Rakhlin · 2018
Later among the works it cites.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Later among the works it cites.
A stochastic line search method with convergence rate analysis
Original
Courtney Paquette and Katya Scheinberg · 2018
Later among the works it cites.
L4: practical loss-based stepsize adaptation for deep learning
Michal Rolinek and Georg Martius · 2018
Later among the works it cites.
Guaranteed sufficient decrease for stochastic variance reduced gradient optimization
Fanhua Shang, Yuanyuan Liu, Kaiwen Zhou, James Cheng, Kelvin Ng, and Yuichi Yoshida · 2018
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason Lee · 2018
Later among the works it cites.
Backtracking gradient descent method for general C 1
Original
Tuyen Trung Truong and Tuan Hang Nguyen · 2018
Later among the works it cites.
Does data interpolation contradict statistical optimality?
Mikhail Belkin, Alexander Rakhlin, and Alexandre B. Tsybakov · 2019
Closest in time.
Training neural networks for and by interpolation
Original
Leonard Berrada, Andrew Zisserman, and M Pawan Kumar · 2019
Closest in time.
Convergence rate analysis of a stochastic trust region method via supermartingales
Jose Blanchet, Coralia Cartis, Matt Menickelly, and Katya Scheinberg · 2019
Closest in time.
Reducing noise in GAN training with variance reduced extragradient
Tatjana Chavdarova, Gauthier Gidel, François Fleuret, and Simon Lacoste-Julien · 2019
Closest in time.
On the ineffectiveness of variance reduced optimization for deep learning
Aaron Defazio and Léon Bottou · 2019
Closest in time.
A variational inequality perspective on generative adversarial networks
Gauthier Gidel, Hugo Berard, Gaëtan Vignoud, Pascal Vincent, and Simon Lacoste-Julien · 2019
Closest in time.
Variance-based extragradient methods with line search for stochastic variational inequalities
Alfredo N Iusem, Alejandro Jofré, Roberto I Oliveira, and Philip Thompson · 2019
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Closest in time.
Accelerating stochastic training for over-parametrized learning
Original
Chaoyue Liu and Mikhail Belkin · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Closest in time.
On the convergence of Adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Closest in time.
Fast and faster convergence of SGD for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Closest in time.