Fetching the paper…
Reading the bibliography…
Reliant on too many experiments to learn good actions, current Reinforcement Learning (RL) algorithms have limited applicability in real-world settings, which can be too expensive to allow exploration.
Individual comparisons by ranking methods
F. Wilcoxon · 1945
Earlier work this paper cites.
Individual comparisons by ranking methods
F. Wilcoxon · 1945
Earlier work this paper cites.
Dynamic Programming
R. E. Bellman · 1957
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
D. A. Pomerleau · 1991
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
S. Thrun and A. Schwartz · 1993
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
S. Thrun and A. Schwartz · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Weak Convergence and Empirical Processes: With Applications to Statistics
A. van der Vaart, A. van der Vaart, A. W. van der Vaart, and J. Wellner · 1996
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
A. Müller · 1997
Earlier work this paper cites.
An overview of statistical learning theory
V. N. Vapnik · 1999
Earlier work this paper cites.
Actor-critic algorithms
V. Konda and J. Tsitsiklis · 2000
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
S. A. Geer and S. van de Geer · 2000
Earlier work this paper cites.
Efficient approximate planning in continuous space markovian decision problems
C. Szepesvári · 2001
Earlier work this paper cites.
Efficient approximate planning in continuous space markovian decision problems
C. Szepesvári · 2001
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Stochastic optimal control: the discrete-time case
D. P. Bertsekas and S. Shreve · 2004
Earlier work this paper cites.
Information theory and statistics: A tutorial
I. Csiszár and P. Shields · 2004
Earlier work this paper cites.
Stochastic optimal control: the discrete-time case
D. P. Bertsekas and S. Shreve · 2004
Earlier work this paper cites.
Unifying divergence minimization and statistical inference via convex duality
Y. Altun and A. Smola · 2006
Earlier work this paper cites.
Fitted Q-iteration in continuous action-space MDPs
A. Antos, R. Munos, and C. Szepesvari · 2007
Earlier work this paper cites.
Fitted Q-iteration in continuous action-space MDPs
A. Antos, R. Munos, and C. Szepesvari · 2007
Earlier work this paper cites.
Reinforcement learning and dynamic programming using function approximators , volume 39
L. Busoniu, R. Babuska, B. De Schutter, and D. Ernst · 2010
Earlier work this paper cites.
Double Q-learning
H. V. Hasselt · 2010
Earlier work this paper cites.
Reinforcement learning and dynamic programming using function approximators , volume 39
L. Busoniu, R. Babuska, B. De Schutter, and D. Ernst · 2010
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
X. Nguyen, M. J. Wainwright, and M. I. Jordan · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Information Theory: Coding Theorems for Discrete Memoryless Systems
I. Csiszár and J. Körner · 2011
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
High confidence policy improvement
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Behavior Regularized Offline Reinforcement Learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Later among the works it cites.
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Maxmin q-learning: Controlling the estimation bias of q-learning
Q. Lan, Y. Pan, A. Fyshe, and M. White · 2019
Later among the works it cites.
Batch policy learning under constraints
H. Le, C. Voloshin, and Y. Yue · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
H. v. Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas · 2016
Cited alongside, same era.
Ucb exploration via q-ensembles
R. Y. Chen, S. Sidor, P. Abbeel, and J. Schulman · 2017
Cited alongside, same era.
Learning from demonstrations for real world reinforcement learning
T. Hester, M. Vecerík, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, A. Sendonaris, G. Dulac-Arnold, I. Osband, J. P. Agapiou, J. Z. Leibo, and A. Gruslys · 2017
Cited alongside, same era.
Y. Wu, G. Tucker, and O. Nachum · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2020
Later among the works it cites.
Ddpg++: Striving for simplicity in continuous-control off-policy reinforcement learning
R. Fakoor, P. Chaudhari, and A. J. Smola · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Interpretable off-policy evaluation in reinforcement learning by highlighting influential transitions
O. Gottesman, J. Futoma, Y. Liu, S. Parbhoo, L. Celi, E. Brunskill, and F. Doshi-Velez · 2020
Later among the works it cites.
Is pessimism provably efficient for offline RL?
Y. Jin, Z. Yang, and Z. Wang · 2020
Later among the works it cites.
Morel : Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Conservative Q-Learning for Offline Reinforcement Learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
A. Kuznetsov, P. Shvechikov, A. Grishin, and D. Vetrov · 2020
Later among the works it cites.
Maxmin q-learning: Controlling the estimation bias of q-learning
Q. Lan, Y. Pan, A. Fyshe, and M. White · 2020
Later among the works it cites.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Provably good batch off-policy reinforcement learning without great exploration
Y. Liu, A. Swaminathan, A. Agarwal, and E. Brunskill · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
T. L. Paine, C. Paduraru, A. Michi, C. Gulcehre, K. Zolna, A. Novikov, Z. Wang, and N. de Freitas · 2020
Later among the works it cites.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
N. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, N. Heess, and M. Riedmiller · 2020
Later among the works it cites.
Z. Wang, A. Novikov, K. Zolna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, and N. de Freitas · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Conservative Q-Learning for Offline Reinforcement Learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
The importance of pessimism in fixed-dataset policy optimization
J. Buckman, C. Gelada, and M. G. Bellemare · 2021
Closest in time.
Benchmarks for deep off-policy evaluation, 2021
J. Fu, M. Norouzi, O. Nachum, G. Tucker, Z. Wang, A. Novikov, M. Yang, M. R. Zhang, Y. Chen, A. Kumar, C. Paduraru, S. Levine, and T. L. Paine · 2021
Closest in time.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
S. K. S. Ghasemipour, D. Schuurmans, and S. S. Gu · 2021
Closest in time.
Regularized behavior value estimation, 2021
C. Gulcehre, S. G. Colmenarejo, Z. Wang, J. Sygnowski, T. Paine, K. Zolna, Y. Chen, M. Hoffman, R. Pascanu, and N. de Freitas · 2021
Closest in time.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
S. K. S. Ghasemipour, D. Schuurmans, and S. S. Gu · 2021
Closest in time.
Regularized behavior value estimation, 2021
C. Gulcehre, S. G. Colmenarejo, Z. Wang, J. Sygnowski, T. Paine, K. Zolna, Y. Chen, M. Hoffman, R. Pascanu, and N. de Freitas · 2021
Closest in time.