Fetching the paper…
Reading the bibliography…
Regardless of the particular task we want them to perform in an environment, there are often shared safety constraints we want our agents to respect.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Andrew Y Ng, Adam Coates, Mark Diel, Varun Ganapathi, Jamie Schulte, Ben Tse, Eric Berger, and Eric Liang · 2006
Earlier work this paper cites.
A control architecture for quadruped locomotion over rough terrain
J Zico Kolter, Mike P Rodgers, and Andrew Y Ng · 2008
Earlier work this paper cites.
Learning to search: Functional gradient techniques for imitation learning
Nathan D Ratliff, David Silver, and J Andrew Bagnell · 2009
Earlier work this paper cites.
On integral probability metrics, \ \backslash phi-divergences and binary classification
Bharath K Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Gert RG Lanckriet · 2009
Earlier work this paper cites.
Learning from demonstration for autonomous navigation in complex unstructured terrain
David Silver, J Andrew Bagnell, and Anthony Stentz · 2010
Earlier work this paper cites.
Follow-the-regularized-leader and mirror descent: Equivalence theorems and l1 regularization
Brendan McMahan · 2011
Earlier work this paper cites.
Optimization and learning for rough terrain legged locomotion
Matt Zucker, Nathan Ratliff, Martin Stolle, Joel Chestnutt, J Andrew Bagnell, Christopher G Atkeson, and James Kuffner · 2011
Earlier work this paper cites.
Activity forecasting
Kris M Kitani, Brian D Ziebart, James Andrew Bagnell, and Martial Hebert · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Probabilistic pointing target prediction via inverse optimal control
Brian Ziebart, Anind Dey, and J Andrew Bagnell · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
Erwin Coumans and Yunfei Bai · 2016
Earlier work this paper cites.
Generative adversarial imitation learning, 2016
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Repeated inverse reinforcement learning
Kareem Amin, Nan Jiang, and Satinder Singh · 2017
Cited alongside, same era.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Multi-task maximum entropy inverse reinforcement learning
Adam Gleave and Oliver Habryka · 2018
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Later among the works it cites.
Projection-based constrained policy optimization
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, and Peter J Ramadge · 2020
Later among the works it cites.
Maximum likelihood constraint inference from stochastic demonstrations
David L McPherson, Kaylene C Stocking, and S Shankar Sastry · 2021
Later among the works it cites.
Of moments and matching: A game-theoretic framework for closing the imitation gap, 2021
Gokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, and Zhiwei Steven Wu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Qingkai Liang, Fanyu Que, and Eytan Modiano · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Learning task specifications from demonstrations
Marcell Vazquez-Chanlatte, Susmit Jha, Ashish Tiwari, Mark K Ho, and Sanjit Seshia · 2018
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Cited alongside, same era.
Maximum likelihood constraint inference for inverse reinforcement learning
Dexter RR Scobee and S Shankar Sastry · 2019
Cited alongside, same era.
Later among the works it cites.
A review of safe reinforcement learning: Methods, theory and applications
Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, Yaodong Yang, and Alois Knoll · 2022
Later among the works it cites.
Efficient off-policy safe reinforcement learning using trust region conditional value at risk
Dohyeong Kim and Songhwai Oh · 2022
Later among the works it cites.
Constrained variational policy optimization for safe reinforcement learning
Zuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu, Steven Wu, Bo Li, and Ding Zhao · 2022
Later among the works it cites.
Minimax optimal online imitation learning via replay estimation
Gokul Swamy, Nived Rajaraman, Matthew Peng, Sanjiban Choudhury, J Andrew Bagnell, Zhiwei Steven Wu, Jiantao Jiao, and Kannan Ramchandran · 2022
Later among the works it cites.
Tianshou: A highly modularized deep reinforcement learning library
Jiayi Weng, Huayu Chen, Dong Yan, Kaichao You, Alexis Duburcq, Minghao Zhang, Yi Su, Hang Su, and Jun Zhu · 2022
Later among the works it cites.
Mengdi Xu, Zuxin Liu, Peide Huang, Wenhao Ding, Zhepeng Cen, Bo Li, and Ding Zhao · 2022
Later among the works it cites.
Learning safety constraints from demonstrations with unknown rewards
David Lindner, Xin Chen, Sebastian Tschiatschek, Katja Hofmann, and Andreas Krause · 2023
Closest in time.
On the robustness of safe reinforcement learning under observational perturbations
Zuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang, Jie Tan, Bo Li, and Ding Zhao · 2023
Closest in time.
Ted Moskovitz, Brendan O’Donoghue, Vivek Veeriah, Sebastian Flennerhag, Satinder Singh, and Tom Zahavy · 2023
Closest in time.
Inverse reinforcement learning without reinforcement learning
Gokul Swamy, Sanjiban Choudhury, J Andrew Bagnell, and Zhiwei Steven Wu · 2023
Closest in time.