Fetching the paper…
Reading the bibliography…
Methods that extract policy primitives from offline demonstrations using deep generative models have shown promise at accelerating reinforcement learning(RL) for new tasks.
Chance-constrained programming
A. Charnes and W. W. Cooper · 1959
Earlier work this paper cites.
The maxq method for hierarchical reinforcement learning
T. G. Dietterich · 1998
Earlier work this paper cites.
Practical reinforcement learning in continuous spaces
W. D. Smart and L. P. Kaelbling · 2000
Earlier work this paper cites.
Integrating guidance into relational reinforcement learning
K. Driessens and S. Dzeroski · 2004
Earlier work this paper cites.
A comparison of tight generalization error bounds
M. Kääriäinen and J. Langford · 2005
Earlier work this paper cites.
Convex approximations of chance constrained programs
A. Nemirovski and A. Shapiro · 2007
Earlier work this paper cites.
Skill discovery in continuous reinforcement learning domains using skill chaining
G. Konidaris and A. Barto · 2009
Earlier work this paper cites.
Safe reinforcement learning with natural language constraints
T. Yang, M. Hu, Y. Chow, P. J. Ramadge, and K. Narasimhan · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
G. Dulac-Arnold, R. Evans, H. van Hasselt, P. Sunehag, T. Lillicrap, J. Hunt, T. Mann, T. Weber, T. Degris, and B. Coppin · 2015
Earlier work this paper cites.
Convex synthesis of randomized policies for controlled markov chains with density safety upper bound constraints
M. El Chamie, Y. Yu, and B. Açıkmeşe · 2016
Earlier work this paper cites.
Constrained policy optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause · 2017
Earlier work this paper cites.
Multi-modal imitation learning from unstructured demonstrations using generative adversarial nets
K. Hausman, Y. Chebotar, S. Schaal, G. Sukhatme, and J. J. Lim · 2017
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Earlier work this paper cites.
Multi-level discovery of deep options
R. Fox, S. Krishnan, I. Stoica, and K. Goldberg · 2017
Earlier work this paper cites.
Density estimation using real NVP
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Multi-level discovery of deep options
R. Fox, S. Krishnan, I. Stoica, and K. Goldberg · 2017
Cited alongside, same era.
Do deep generative models know what they don’t know?
E. Nalisnick, A. Matsukawa, Y. Teh, D. Gorur, and B. Lakshminarayanan · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
G. Dalal, K. Dvijotham, M. Vecerík, T. Hester, C. Paduraru, and Y. Tassa · 2018
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Y. Chow, O. Nachum, E. Duenez-Guzman, and M. Ghavamzadeh · 2018
Cited alongside, same era.
Learning model predictive control for iterative tasks. a data-driven control framework
U. Rosolia and F. Borrelli · 2018
Cited alongside, same era.
Safe reinforcement learning in constrained Markov decision processes
A. Wachi and Y. Sui · 2020
Later among the works it cites.
Projection-based constrained policy optimization
K. Narasimhan · 2020
Later among the works it cites.
Safe reinforcement learning via curriculum induction
M. Turchetta, A. Kolobov, S. Shah, A. Krause, and A. Agarwal · 2020
Later among the works it cites.
Learning to be safe: Deep rl with a safety critic
K. P. Srinivasan, B. Eysenbach, S. Ha, J. Tan, and C. Finn · 2020
Later among the works it cites.
Safety augmented value estimation from demonstrations (saved): Safe deep model-based rl for sparse cost robotic tasks
B. Thananjeyan, A. Balakrishna, U. Rosolia, F. Li, R. McAllister, J. Gonzalez, S. Levine, F. Borrelli, and K. Goldberg · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data-efficient hierarchical reinforcement learning
O. Nachum, S. S. Gu, H. Lee, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
J. Achiam and D. Amodei · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2019
Cited alongside, same era.
Near-optimal representation learning for hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2019
Cited alongside, same era.
MCP: Learning Composable Hierarchical Control with Multiplicative Compositional Policies
X. B. Peng, M. Chang, G. Zhang, P. Abbeel, and S. Levine · 2019
Cited alongside, same era.
Learning action representations for reinforcement learning
Y. Chandak, G. Theocharous, J. Kostas, S. Jordan, and P. Thomas · 2019
Cited alongside, same era.
Dynamics-aware unsupervised discovery of skills
A. Sharma, S. S. Gu, S. Levine, V. Kumar, and K. Hausman · 2020
Later among the works it cites.
Learning robot skills with temporal variational inference
T. Shankar and A. Gupta · 2020
Later among the works it cites.
Data-efficient visuomotor policy training using reinforcement learning and generative models
A. Ghadirzadeh, P. Poklukar, V. Kyrki, D. Kragic, and M. Björkman · 2020
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
A. Singh, H. Liu, G. Zhou, A. Yu, N. Rhinehart, and S. Levine · 2021
Later among the works it cites.
Guided reinforcement learning with learned skills
K. Pertsch, Y. Lee, Y. Wu, and J. J. Lim · 2021
Later among the works it cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2021
Later among the works it cites.
Conservative safety critics for exploration
H. Bharadhwaj, A. Kumar, N. Rhinehart, S. Levine, F. Shkurti, and A. Garg · 2021
Later among the works it cites.
Accelerating safe reinforcement learning with constraint-mismatched policies
T.-Y. Yang, J. P. Rosca, K. Narasimhan, and P. J. Ramadge · 2021
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
B. Thananjeyan, A. Balakrishna, S. Nair, M. Luo, K. P. Srinivasan, M. Hwang, J. E. Gonzalez, J. Ibarz, C. Finn, and K. Goldberg · 2021
Later among the works it cites.
Abc-lmpc: Safe sample-based learning mpc for stochastic nonlinear dynamical systems with adjustable boundary conditions
B. Thananjeyan, A. Balakrishna, U. Rosolia, J. Gonzalez, A. D. Ames, and K. Goldberg · 2021
Later among the works it cites.
Latent skill planning for exploration and transfer
K. Xie, H. Bharadhwaj, D. Hafner, A. Garg, and F. Shkurti · 2021
Later among the works it cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
E. Coumans and Y. Bai · 2021
Later among the works it cites.