Fetching the paper…
Reading the bibliography…
In recent years, deep off-policy actor-critic algorithms have become a dominant approach to reinforcement learning for continuous control.
Robust estimation of a location parameter
P. J. Huber · 1964
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L.-J. Lin · 1992
Earlier work this paper cites.
Q-learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
S. Thrun and A. Schwartz · 1993
Earlier work this paper cites.
Model selection in contextual stochastic bandit problems
A. Pacchiano, M. Phan, Y. Abbasi-Yadkori, A. Rao, J. Zimmert, T. Lattimore, and C. Szepesvari · 2003
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2004
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Bandit based Monte-Carlo planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Tuning bandit algorithms in stochastic environments
J. Audibert, R. Munos, and C. Szepesvári · 2007
Earlier work this paper cites.
Optimism in reinforcement learning and kullback-leibler divergence
S. Filippi, O. Cappé, and A. Garivier · 2010
Earlier work this paper cites.
Double q-learning
H. Hasselt · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
P. L. Bartlett and A. Tewari · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2012
Earlier work this paper cites.
Regret bound balancing and elimination for model selection in bandits and rl
A. Pacchiano, C. Dann, C. Gentile, and P. Bartlett · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Openai gym
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
Corralling a band of bandit algorithms
A. Agarwal, H. Luo, B. Neyshabur, and R. E. Schapire · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Statistics and samples in distributional reinforcement learning, 2019
M. Rowland, R. Dadashi, S. Kumar, R. Munos, M. G. Bellemare, and W. Dabney · 2019
Later among the works it cites.
Adapting behaviour for learning progress
T. Schaul, D. Borsa, D. Ding, D. Szepesvari, G. Ostrovski, W. Dabney, and S. Osindero · 2019
Later among the works it cites.
Near-optimal optimistic reinforcement learning using empirical bernstein inequalities
A. C. Y. Tossou, D. Basu, and C. Dimitrakakis · 2019
Later among the works it cites.
Quota: The quantile option architecture for reinforcement learning
S. Zhang and H. Yao · 2019
Later among the works it cites.
Agent57: Outperforming the atari human benchmark
A. P. Badia, B. Piot, S. Kapturowski, P. Sprechmann, A. Vitvitskyi, D. Guo, and C. Blundell · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributional policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. TB, A. Muldal, N. Heess, and T. Lillicrap · 2018
Cited alongside, same era.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
R. Fruit, M. Pirotta, A. Lazaric, and R. Ortner · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Ready policy one: World building through active learning
P. Ball, J. Parker-Holder, A. Pacchiano, K. Choromanski, and S. Roberts · 2020
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
C. Jin, Z. Yang, Z. Wang, and M. I. Jordan · 2020
Later among the works it cites.
Improving generalization in meta reinforcement learning using learned objectives
L. Kirsch, S. van Steenkiste, and J. Schmidhuber · 2020
Later among the works it cites.
Reinforcement learning with augmented data
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas · 2020
Later among the works it cites.
Predictive information accelerates learning in rl
K.-H. Lee, I. Fischer, A. Liu, Y. Guo, H. Lee, J. Canny, and S. Guadarrama · 2020
Later among the works it cites.
Discovering reinforcement learning algorithms
J. Oh, M. Hessel, W. M. Czarnecki, Z. Xu, H. van Hasselt, S. Singh, and D. Silver · 2020
Later among the works it cites.
Effective diversity in population-based reinforcement learning
J. Parker-Holder, A. Pacchiano, K. Choromanski, and S. Roberts · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
L. Yang and M. Wang · 2020
Later among the works it cites.
Offcon 3 : What is state of the art anyway?, 2021
P. J. Ball and S. J. Roberts · 2021
Closest in time.
Evolving reinforcement learning algorithms
J. D. Co-Reyes, Y. Miao, D. Peng, Q. V. Le, S. Levine, H. Lee, and A. Faust · 2021
Closest in time.
Efficient wasserstein natural gradients for reinforcement learning
T. Moskovitz, M. Arbel, F. Huszar, and A. Gretton · 2021
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
D. Yarats, I. Kostrikov, and R. Fergus · 2021
Closest in time.