Fetching the paper…
Reading the bibliography…
Learning policy from offline datasets through offline reinforcement learning (RL) holds promise for scaling data-driven decision-making while avoiding unsafe and costly online interactions.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020) · 2004
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
Li, L., Yang, R., and Luo, D. (2020) · 2010
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (2017) · 2017
Earlier work this paper cites.
Regularizing and optimizing lstm language models
Merity, S., Keskar, N. S., and Socher, R. (2017) · 2017
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Wang, Q., Xiong, J., Han, L., Liu, H., Zhang, T., et al. (2018) · 2018
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020) · 2020
Earlier work this paper cites.
Behavioral cloning from noisy demonstrations
Sasaki, F. and Yamashina, R. (2020) · 2020
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T. (2020) · 2020
Earlier work this paper cites.
Robust deep reinforcement learning against adversarial perturbations on state observations
Zhang, H., Chen, H., Xiao, C., Li, B., Liu, M., Boning, D., and Hsieh, C.-J. (2020) · 2020
Earlier work this paper cites.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
An, G., Moon, S., Kim, J.-H., and Song, H. O. (2021) · 2021
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021) · 2021
Earlier work this paper cites.
Rvs: What is essential for offline rl via supervised learning?
Emmons, S., Eysenbach, B., Kostrikov, I., and Levine, S. (2021) · 2021
Earlier work this paper cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S. (2021) · 2021
Earlier work this paper cites.
Generalized decision transformer for offline hindsight information matching
Furuta, H., Matsuo, Y., and Gu, S. S. (2021) · 2021
Earlier work this paper cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S. (2021) · 2021
Earlier work this paper cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S. (2021) · 2021
Earlier work this paper cites.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Bai, C., Wang, L., Yang, Z., Deng, Z., Garg, A., Liu, P., and Wang, Z. (2022) · 2022
Cited alongside, same era.
Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters
Ghasemipour, K., Gu, S. S., and Nachum, O. (2022) · 2022
Cited alongside, same era.
Decision transformer under random frame dropping
Hu, K., Zheng, R. C., Gao, Y., and Xu, H. (2022) · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J., and Levine, S. (2022) · 2022
Cited alongside, same era.
Robust reinforcement learning using offline data
Panaganti, K., Xu, Z., Kalathil, D., and Ghavamzadeh, M. (2022) · 2022
Cited alongside, same era.
Emergent agentic transformer from chain of hindsight experience
Liu, H. and Abbeel, P. (2023) · 2023
Later among the works it cites.
Anti-exploration by random network distillation
Nikulin, A., Kurenkov, V., Tarasov, D., and Kolesnikov, S. (2023) · 2023
Later among the works it cites.
Unleashing the power of pre-trained language models for offline reinforcement learning
Shi, R., Liu, Y., Ze, Y., Du, S. S., and Xu, H. (2023) · 2023
Later among the works it cites.
Goplan: Goal-conditioned offline reinforcement learning by planning with learned models
Wang, M., Yang, R., Chen, X., and Fang, M. (2023) · 2023
Later among the works it cites.
Future-conditioned unsupervised pretraining for decision transformer
Xie, Z., Lin, Z., Ye, D., Fu, Q., Wei, Y., and Li, S. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shi, L. and Chi, Y. (2022) · 2022
Cited alongside, same era.
Exploit reward shifting in value-based deep-rl: Optimistic curiosity-based exploration and conservative exploitation via linear reward shaping
Sun, H., Han, L., Yang, R., Ma, X., Guo, J., and Zhou, B. (2022) · 2022
Cited alongside, same era.
Dasco: Dual-generator adversarial support constrained offline reinforcement learning
Vuong, Q., Kumar, A., Levine, S., and Chebotar, Y. (2022) · 2022
Cited alongside, same era.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M. (2022) · 2022
Cited alongside, same era.
Copa: Certifying robust policies for offline reinforcement learning against poisoning attacks
Wu, F., Li, L., Xu, C., Zhang, H., Kailkhura, B., Kenthapadi, K., Zhao, D., and Li, B. (2022) · 2022
Cited alongside, same era.
Prompting decision transformer for few-shot policy generalization
Xu, M., Shen, Y., Zhang, S., Lu, Y., Zhao, D., Tenenbaum, J., and Gan, C. (2022) · 2022
Cited alongside, same era.
Corruption-robust offline reinforcement learning
Zhang, X., Chen, Y., Zhu, X., and Sun, W. (2022) · 2022
Cited alongside, same era.
Offline rl with no ood actions: In-sample learning via implicit value regularization
Xu, H., Jiang, L., Li, J., Yang, Z., Wang, Z., Chan, V. W. K., and Zhan, X. (2023) · 2023
Later among the works it cites.
Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Yamagata, T., Khalil, A., and Santos-Rodriguez, R. (2023) · 2023
Later among the works it cites.
Zhihe, Y. and Xu, Y. (2023) · 2023
Later among the works it cites.
Exact policy recovery in offline rl with both heavy-tailed rewards and data corruption
Chen, Y., Zhang, X., Xie, Q., and Zhu, X. (2024) · 2024
Closest in time.
Offline reinforcement learning with behavior value regularization
Huang, L., Dong, B., Xie, W., and Zhang, W. (2024) · 2024
Closest in time.
Ropo: Robust preference optimization for large language models
Liang, X., Chen, C., Qiu, S., Wang, J., Wu, Y., Fu, Z., Shi, Z., Wu, F., and Ye, J. (2024) · 2024
Closest in time.
Corruption robust offline reinforcement learning with human feedback
Mandal, D., Nika, A., Kamalaruban, P., Singla, A., and Radanović, G. (2024) · 2024
Closest in time.
Hiql: Offline goal-conditioned rl with latent states as actions
Park, S., Ghosh, D., Eysenbach, B., and Levine, S. (2024) · 2024
Closest in time.
Accountability in offline reinforcement learning: Explaining decisions with a corpus of examples
Sun, H., Hüyük, A., Jarrett, D., and van der Schaar, M. (2024) · 2024
Closest in time.
Elastic decision transformer
Wu, Y.-H., Wang, X., and Hamaya, M. (2024) · 2024
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D. (2019) · 2062
Closest in time.