Fetching the paper…
Reading the bibliography…
Safe reinforcement learning (RL) trains a constraint satisfaction policy by interacting with the environment.
Algaedice: Policy gradient from arbitrary experience
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D · 1912
Earlier work this paper cites.
Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Altman, E · 1998
Earlier work this paper cites.
Gendice: Generalized offline estimation of stationary values
Zhang, R., Dai, B., Li, L., and Schuurmans, D · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M · 2017
Earlier work this paper cites.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W., Sastry, G., Stuhlmueller, A., and Evans, O · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Earlier work this paper cites.
Lyapunov-based safe policy optimization for continuous control
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Batch policy learning under constraints
Le, H., Voloshin, C., and Yue, Y · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D · 2019
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Gulcehre, C., Wang, Z., Novikov, A., Paine, T., Gómez, S., Zolna, K., Agarwal, R., Merel, J. S., Mankowitz, D. J., Paduraru, C., et al · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Xie, T., Jiang, N., Wang, H., Xiong, C., and Bai, Y · 2021
Later among the works it cites.
Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning
Yang, Q., Simão, T. D., Tindemans, S. H., and Spaan, M. T · 2021
Later among the works it cites.
Model-free safe control for zero-violation reinforcement learning
Zhao, W., He, T., and Liu, C · 2021
Later among the works it cites.
Saac: Safe reinforcement learning as an adversarial game of actor-critics
Flet-Berliac, Y. and Basu, D · 2022
Later among the works it cites.
Generalized decision transformer for offline hindsight information matching
Furuta, H., Matsuo, Y., and Gu, S. S · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Constrained model-based reinforcement learning with robust cross-entropy method
Liu, Z., Zhou, H., Chen, B., Zhong, S., Hebert, M., and Zhao, D · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Cited alongside, same era.
Responsive safety in reinforcement learning by pid lagrangian methods
Stooke, A., Achiam, J., and Abbeel, P · 2020
Cited alongside, same era.
Scalability in perception for autonomous driving: Waymo open dataset
Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., et al · 2020
Cited alongside, same era.
Critic regularized regression
Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., Heess, N., et al · 2020
Cited alongside, same era.
Projection-based constrained policy optimization
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J · 2020
Cited alongside, same era.
Gong, C., Yang, Z., Bai, Y., He, J., Shi, J., Sinha, A., Xu, B., Hou, X., Fan, G., and Lo, D · 2022
Later among the works it cites.
Bullet-safety-gym: Aframework for constrained reinforcement learning
Gronauer, S · 2022
Later among the works it cites.
A review of safe reinforcement learning: Methods, theory and applications
Gu, S., Yang, L., Du, Y., Chen, G., Walter, F., Wang, J., Yang, Y., and Knoll, A · 2022
Later among the works it cites.
Lee, J., Paduraru, C., Mankowitz, D. J., Heess, N., Precup, D., Kim, K.-E., and Guez, A · 2022
Later among the works it cites.
Lu, Y., Fu, J., Tucker, G., Pan, X., Bronstein, E., Roelofs, B., Sapp, B., White, B., Faust, A., Whiteson, S., et al · 2022
Later among the works it cites.
Constrained offline policy optimization
Polosky, N., Da Silva, B. C., Fiterau, M., and Jagannath, J · 2022
Later among the works it cites.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Prudencio, R. F., Maximo, M. R., and Colombini, E. L · 2022
Later among the works it cites.
Offline reinforcement learning from heteroskedastic data via support constraints
Singh, A., Kumar, A., Chebotar, Y., Levine, S., et al · 2022
Later among the works it cites.
S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics
Sinha, S., Mandlekar, A., and Garg, A · 2022
Later among the works it cites.
Sauté rl: Almost surely safe reinforcement learning using state augmentation
Sootla, A., Cowen-Rivers, A. I., Jafferjee, T., Wang, Z., Mguni, D. H., Wang, J., and Ammar, H · 2022
Later among the works it cites.
CORL: Research-oriented deep offline reinforcement learning library
Tarasov, D., Nikulin, A., Akimov, D., Kurenkov, V., and Kolesnikov, S · 2022
Later among the works it cites.
Zheng, Q., Zhang, A., and Grover, A · 2022
Later among the works it cites.
Omnisafe: An infrastructure for accelerating safe reinforcement learning research
Ji, J., Zhou, J., Zhang, B., Dai, J., Pan, X., Sun, R., Huang, W., Geng, Y., Liu, M., and Yang, Y · 2023
Closest in time.
Datasets and benchmarks for offline safe reinforcement learning
Liu, Z., Guo, Z., Lin, H., Yao, Y., Zhu, J., Cen, Z., Hu, H., Yu, W., Zhang, T., Tan, J., et al · 2023
Closest in time.