Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) in low-data and risk-sensitive domains requires performant and flexible deployment policies that can readily incorporate constraints during deployment.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
S. Thrun and A. Schwartz · 1993
Earlier work this paper cites.
An upper bound on the loss from approximate optimal-value functions
S. P. Singh and R. C. Yee · 1994
Earlier work this paper cites.
Constrained markov decision processes with total cost criteria: Occupation measures and primal lp
E. Altman · 1996
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
E. Altman · 1999
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
R. Rubinstein · 1999
Earlier work this paper cites.
Error bounds for approximate value iteration
R. Munos · 2005
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
Natural actor-critic
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvári · 2008
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mulling, and Y. Altun · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Guided policy search
S. Levine and V. Koltun · 2013
Earlier work this paper cites.
Model predictive control
E. F. Camacho and C. B. Alba · 2013
Earlier work this paper cites.
Aggressive driving with model predictive path integral control
G. Williams, P. Drews, B. Goldfain, J. M. Rehg, and E. A. Theodorou · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2016
Earlier work this paper cites.
Constrained policy optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Earlier work this paper cites.
Learning to walk via deep reinforcement learning
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine · 2018
Earlier work this paper cites.
Sim-to-real reinforcement learning for deformable object manipulation
J. Matas, S. James, and A. J. Davison · 2018
Earlier work this paper cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Van Hoof, and D. Meger · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2018
Cited alongside, same era.
Non-delusional q-learning and value-iteration
T. Lu, D. Schuurmans, and C. Boutilier · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
V. Feinberg, A. Wan, I. Stoica, M. I. Jordan, J. E. Gonzalez, and S. Levine · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Cited alongside, same era.
Online planning with lookahead policies
Y. Efroni, M. Ghavamzadeh, and S. Mannor · 2020
Closest in time.
Model-augmented actor-critic: Backpropagating through paths
I. Clavera, V. Fu, and P. Abbeel · 2020
Closest in time.
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2020
Closest in time.
Gendice: Generalized offline estimation of stationary values
R. Zhang, B. Dai, L. Li, and D. Schuurmans · 2020
Closest in time.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
N. Y. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, and M. Riedmiller · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Cited alongside, same era.
Probabilistic planning with sequential monte carlo methods
A. Piché, V. Thomas, C. Ibrahim, Y. Bengio, and C. Pal · 2018
Cited alongside, same era.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Y. Luo, H. Xu, Y. Li, Y. Tian, T. Darrell, and T. Ma · 2018
Cited alongside, same era.
Diagnosing bottlenecks in deep q-learning algorithms
J. Fu, A. Kumar, M. Soh, and S. Levine · 2019
Cited alongside, same era.
Towards characterizing divergence in deep q-learning
J. Achiam, E. Knight, and P. Abbeel · 2019
Cited alongside, same era.
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Closest in time.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Closest in time.
Plas: Latent action space for offline reinforcement learning
W. Zhou, S. Bajracharya, and D. Held · 2020
Closest in time.
A. Argenson and G. Dulac-Arnold · 2020
Closest in time.
Deployment-efficient reinforcement learning via model-based offline optimization
T. Matsushima, H. Furuta, Y. Matsuo, O. Nachum, and S. Gu · 2020
Closest in time.
Leverage the average: an analysis of regularization in rl
N. Vieillard, T. Kozuno, B. Scherrer, O. Pietquin, R. Munos, and M. Geist · 2020
Closest in time.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Closest in time.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Closest in time.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Closest in time.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Closest in time.
Z. Wang, A. Novikov, K. Żołna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, et al · 2020
Closest in time.
Safe model-based reinforcement learning with robust cross-entropy method
Z. Liu, H. Zhou, B. Chen, S. Zhong, M. Hebert, and D. Zhao · 2020
Closest in time.
Constrained cross-entropy method for safe reinforcement learning
M. Wen and U. Topcu · 2020
Closest in time.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Closest in time.
On the importance of hyperparameter optimization for model-based reinforcement learning
B. Zhang, R. Rajan, L. Pineda, N. Lambert, A. Biedenkapp, K. Chua, F. Hutter, and R. Calandra · 2021
Closest in time.
Lyapunov barrier policy optimization
H. Sikchi, W. Zhou, and D. Held · 2021
Closest in time.