Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) is challenged by the distributional shift between learning policies and datasets.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Self-imitation learning
Oh, J., Guo, Y., Singh, S., and Lee, H · 2018
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Wang, Q., Xiong, J., Han, L., Liu, H., Zhang, T., et al · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Àgata Lapedriza, Jones, N., Gu, S., and Picard, R. W · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
The importance of pessimism in fixed-dataset policy optimization
Buckman, J., Gelada, C., and Bellemare, M. G · 2020
Cited alongside, same era.
BAIL: best-action imitation learning for batch deep reinforcement learning
Chen, X., Zhou, Z., Wang, Z., Wang, C., Wu, Y., and Ross, K · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Later among the works it cites.
Critic regularized regression
Wang, Z., Novikov, A., Zolna, K., Merel, J., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N. Y., Gülçehre, Ç., Heess, N., and de Freitas, N · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Gulcehre, C., Wang, Z., Novikov, A., Paine, T., Gómez, S., Zolna, K., Agarwal, R., Merel, J. S., Mankowitz, D. J., Paduraru, C., et al · 2020
Cited alongside, same era.
Decoupling representation and classifier for long-tailed recognition
Kang, B., Xie, S., Rohrbach, M., Yan, Z., Gordo, A., Feng, J., and Kalantidis, Y · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Cited alongside, same era.
Offline reinforcement learning with fisher divergence critic regularization
Kostrikov, I., Fergus, R., Tompson, J., and Nachum, O
Cited in the paper.
Kostrikov, I., Nair, A., and Levine, S · 2021
Later among the works it cites.
A rank-based sampling framework for offline reinforcement learning
Shen, Y., Fang, Z., Xu, Y., Cao, Y., and Zhu, J · 2021
Later among the works it cites.
Deep long-tailed learning: A survey
Zhang, Y., Kang, B., Hooi, B., Yan, S., and Feng, J · 2021
Later among the works it cites.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
Yarats, D., Brandfonbrener, D., Liu, H., Laskin, M., Abbeel, P., Lazaric, A., and Pinto, L · 2022
Closest in time.