Fetching the paper…
Reading the bibliography…
Model-free reinforcement learning algorithms combined with value function approximation have recently achieved impressive performance in a variety of application domains.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 1905
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Temporal differences-based policy iteration and applications in neuro-dynamic programming
Dimitri P Bertsekas and Sergey Ioffe · 1996
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L. Bartlett · 2009
Earlier work this paper cites.
Online markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Value function approximation in reinforcement learning using the fourier basis
George Konidaris, Sarah Osentoski, and Philip Thomas · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Online learning with predictable sequences
Alexander Rakhlin and Karthik Sridharan · 2012
Earlier work this paper cites.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2012
Earlier work this paper cites.
Online markov decision processes under bandit feedback
G. Neu, A. Gyorgy, C. Szepesvari, and A. Antos · 2013
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Sasha Rakhlin and Karthik Sridharan · 2013
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J Andrew Bagnell · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Accelerating online convex optimization via adaptive prediction
Mehryar Mohri and Scott Yang · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Variance-aware regret bounds for undiscounted reinforcement learning in mdps
Mohammad Sadegh Talebi and Odalric-Ambrym Maillard · 2018
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Later among the works it cites.
Predictor-corrector policy optimization
Ching-An Cheng, Xinyan Yan, Nathan Ratliff, and Byron Boots · 2019
Later among the works it cites.
A theory of regularized Markov decision processes
Mathieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
Exploration bonus for regret minimization in discrete and continuous average reward mdps
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Pooria Joulani, András György, and Csaba Szepesvári · 2017
Cited alongside, same era.
A survey of algorithms and analysis for adaptive online learning
H Brendan McMahan · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Cited alongside, same era.
Deep exploration via randomized value functions
Ian Osband, Daniel Russo, Z Wen, and B Van Roy · 2017
Cited alongside, same era.
Learning unknown markov decision processes: A thompson sampling approach
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar, and Rahul Jain · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
QIAN Jian, Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2019
Later among the works it cites.
Momentum in reinforcement learning
Nino Vieillard, Bruno Scherrer, Olivier Pietquin, and Matthieu Geist · 2019
Later among the works it cites.
Model-free reinforcement learning in infinite-horizon average-reward markov decision processes, 2019
Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, Hiteshi Sharma, and Rahul Jain · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Closest in time.
Leverage the average: an analysis of regularization in rl
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Rémi Munos, and Matthieu Geist · 2020
Closest in time.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Closest in time.