Fetching the paper…
Reading the bibliography…
Offline safe RL is of great practical relevance for deploying agents in real-world applications.
Constrained Markov decision processes: stochastic modeling
Altman, E · 1999
Earlier work this paper cites.
Control and optimization meet the smart power grid: Scheduling of power demands for optimal energy management
Koutsopoulos, I. and Tassiulas, L · 2011
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Deep reinforcement learning framework for autonomous driving
Sallab, A. E., Abdou, M., Perot, E., and Yogamani, S · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Batch policy learning under constraints
Le, H., Voloshin, C., and Yue, Y · 2019
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Critic regularized regression
Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., Heess, N., et al · 2020
Later among the works it cites.
Projection-based constrained policy optimization
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
Coptidice: Offline constrained reinforcement learning via stationary distribution correction estimation
Lee, J., Paduraru, C., Mankowitz, D. J., Heess, N., Precup, D., Kim, K.-E., and Guez, A · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Ipo: Interior-point policy optimization under constraints
Liu, Y., Ding, J., and Liu, X · 2020
Cited alongside, same era.
Constraints penalized q-learning for safe offline reinforcement learning
Xu, H., Zhan, X., and Zhu, X
Cited in the paper.
Prompting decision transformer for few-shot policy generalization
Xu, M., Shen, Y., Zhang, S., Lu, Y., Zhao, D., Tenenbaum, J., and Gan, C
Cited in the paper.
Hu, S., Shen, L., Zhang, Y., Chen, Y., and Tao, D · 2022
Later among the works it cites.
Constrained offline policy optimization
Polosky, N., Da Silva, B. C., Fiterau, M., and Jagannath, J · 2022
Later among the works it cites.
Bootstrapped transformer for offline reinforcement learning
Wang, K., Zhao, H., Luo, X., Ren, K., Zhang, W., and Li, D · 2022
Later among the works it cites.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A · 2022
Later among the works it cites.