Fetching the paper…
Reading the bibliography…
Online learning is a powerful tool for analyzing iterative algorithms.
Mean value methods in iteration
W Robert Mann · 1953
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
Regret bounds for prediction problems
Geoffrey J Gordon · 1999
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Prox-method with rate of convergence o (1/t) for variational inequalities with lipschitz continuous monotone operators and smooth convex-concave saddle point problems
Arkadi Nemirovski · 2004
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Finite-dimensional variational inequalities and complementarity problems
Francisco Facchinei and Jong-Shi Pang · 2007
Earlier work this paper cites.
The complexity of computing a nash equilibrium
Constantinos Daskalakis, Paul W Goldberg, and Christos H Papadimitriou · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Solving variational inequalities with stochastic mirror-prox algorithm
Anatoli Juditsky, Arkadi Nemirovski, and Claire Tauvel · 2011
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2012
Cited alongside, same era.
Revisiting frank-wolfe: Projection-free sparse convex optimization
Martin Jaggi · 2013
Cited alongside, same era.
Non-stationary stochastic optimization
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2015
Cited alongside, same era.
Online optimization: Competing with dynamic comparators
Ali Jadbabaie, Alexander Rakhlin, Shahin Shahrampour, and Karthik Sridharan · 2015
Cited alongside, same era.
Improving multi-step prediction of learned time series models
Arun Venkatraman, Martial Hebert, and J Andrew Bagnell · 2015
Cited alongside, same era.
Improved dynamic regret for non-degenerate functions
Lijun Zhang, Tianbao Yang, Jinfeng Yi, Jing Rong, and Zhi-Hua Zhou · 2017
Later among the works it cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Wen Sun, Arun Venkatraman, Geoffrey J Gordon, Byron Boots, and J Andrew Bagnell · 2017
Later among the works it cites.
Convergence of value aggregation for imitation learning
Ching-An Cheng and Byron Boots · 2018
Later among the works it cites.
A dynamic regret analysis and adaptive regularization algorithm for on-policy robot imitation learning
Jonathan Lee, Michael Laskey, Ajay Kumar Tanwani, Anil Aswani, and Ken Goldberg · 2018
Later among the works it cites.
Online learning with continuous variations: Dynamic regret and reductions
Ching-An Cheng, Jonathan Lee, Ken Goldberg, and Byron Boots · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Introduction to online convex optimization
Elad Hazan et al · 2016
Cited alongside, same era.
Online optimization in dynamic environments: Improved regret rates for strongly convex problems
Aryan Mokhtari, Shahin Shahrampour, Ali Jadbabaie, and Alejandro Ribeiro · 2016
Cited alongside, same era.
Tracking slowly moving clairvoyant: optimal dynamic regret of online learning with true and noisy gradient
Tianbao Yang, Lijun Zhang, Rong Jin, and Jinfeng Yi · 2016
Cited alongside, same era.
Online learning with inexact proximal online gradient descent algorithms
Rishabh Dixit, Amrit Singh Bedi, Ruchi Tripathi, and Ketan Rajawat · 2019
Closest in time.
A reduction from reinforcement learning to no-regret online learning
Ching-An Cheng, Remi Tachet des Combes, Byron Boots, and Geoff Gordon · 2019
Closest in time.
Accelerating imitation learning with predictive models
Ching-An Cheng, Xinyan Yan, Evangelos A Theodorou, and Byron Boots · 2019
Closest in time.
Predictor-corrector policy optimization
Ching-An Cheng, Xinyan Yan, Nathan Ratliff, and Byron Boots · 2019
Closest in time.