Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) is challenged by the distributional shift problem.
J. Tsitsiklis and B. Van Roy, “An analysis of temporal-difference learning with function approximationtechnical,” Rep. LIDS-P-2322). Lab. Inf. Decis. Syst. Massachusetts Inst. Technol. Tech. Rep , 1996
1996
Earlier work this paper cites.
S. Kakade and J. Langford, “Approximately optimal approximate reinforcement learning,” in Proceedings of the Nineteenth International Conference on Machine Learning , 2002, pp. 267–274
2002
Earlier work this paper cites.
H. Hasselt, “Double q-learning,” in Advances in neural information processing systems , vol. 23, 2010
2010
Earlier work this paper cites.
S. Lange, T. Gabel, and M. Riedmiller, “Batch reinforcement learning,” in Reinforcement learning . Springer, 2012
2012
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Oh, Y. Guo, S. Singh, and H. Lee, “Self-imitation learning,” in International conference on machine learning , 2018, pp. 3878–3887
2018
Earlier work this paper cites.
Q. Wang, J. Xiong, L. Han, H. Liu, T. Zhang et al. , “Exponentially weighted imitation learning for batched historical data,” in Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in International conference on machine learning , 2018, pp. 1587–1596
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine, “Stabilizing off-policy q-learning via bootstrapping error reduction,” in Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Kumar, A. Gupta, and S. Levine, “Discor: Corrective feedback in reinforcement learning via distribution correction,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 18 560–18 572
2020
Earlier work this paper cites.
C. Wang, Y. Wu, Q. Vuong, and K. Ross, “Striving for simplicity and performance in off-policy drl: Output normalization and non-uniform sampling,” in International Conference on Machine Learning , 2020, pp. 10 070–10 080
2020
Cited alongside, same era.
X. Chen, Z. Zhou, Z. Wang, C. Wang, Y. Wu, and K. Ross, “Bail: Best-action imitation learning for batch deep reinforcement learning,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 18 353–18 363
2020
Cited alongside, same era.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 1179–1191
2020
Cited alongside, same era.
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine, “D4rl: Datasets for deep data-driven reinforcement learning,” 2020
2020
Cited alongside, same era.
S. Sinha, J. Song, A. Garg, and S. Ermon, “Experience replay with likelihood-free importance weights,” in Learning for Dynamics and Control Conference , 2022
2022
Later among the works it cites.
Y. Yue, B. Kang, X. Ma, Z. Xu, G. Huang, and S. YAN, “Boosting offline reinforcement learning via data rebalancing,” in 3rd NeurIPS Offline RL Workshop’ , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Ghasemipour, S. S. Gu, and O. Nachum, “Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters,” in Advances in Neural Information Processing Systems , vol. 35, 2022, pp. 18 267–18 281
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Wang, A. Novikov, K. Zolna, J. S. Merel, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Siegel, C. Gulcehre, N. Heess et al. , “Critic regularized regression,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 7768–7778
2020
Cited alongside, same era.
2020
Cited alongside, same era.
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learning via sequence modeling,” in Advances in neural information processing systems , vol. 34, 2021, pp. 15 084–15 097
2021
Cited alongside, same era.
S. Fujimoto and S. S. Gu, “A minimalist approach to offline reinforcement learning,” in Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS) , vol. 34, 2021, pp. 20 132–20 145
2021
Cited alongside, same era.
X.-H. Liu, Z. Xue, J.-C. Pang, S. Jiang, F. Xu, and Y. Yu, “Regret minimization experience replay in off-policy reinforcement learning,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
M. Liu, H. Zhao, Z. Yang, J. Shen, W. Zhang, L. Zhao, and T.-Y. Liu, “Curriculum offline imitating learning,” in Advances in Neural Information Processing Systems , vol. 34, 2021, pp. 6266–6277
2021
Cited alongside, same era.
2021
Cited alongside, same era.
D. Brandfonbrener, W. Whitney, R. Ranganath, and J. Bruna, “Offline rl without off-policy evaluation,” in Advances in neural information processing systems , vol. 34, 2021, pp. 4933–4946
2021
Cited alongside, same era.
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin, “Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,” in Conference on Robot Learning , 2022
2022
Later among the works it cites.
V. H. Pong, A. V. Nair, L. M. Smith, C. Huang, and S. Levine, “Offline meta-reinforcement learning with online self-supervision,” in International Conference on Machine Learning , 2022
2022
Later among the works it cites.
Z.-W. Hong, P. Agrawal, R. T. des Combes, and R. Laroche, “Harnessing mixed offline reinforcement learning datasets via trajectory weighting,” in The Eleventh International Conference on Learning Representations , 2023
2023
Closest in time.
R. F. Prudencio, M. R. Maximo, and E. L. Colombini, “A survey on offline reinforcement learning: Taxonomy, review, and open problems,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Closest in time.
A. Singh, A. Kumar, Q. Vuong, Y. Chebotar, and S. Levine, “ReDS: Offline RL with heteroskedastic datasets via support constraints,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Closest in time.
Z. Wang, J. J. Hunt, and M. Zhou, “Diffusion policies as an expressive policy class for offline reinforcement learning,” in The Eleventh International Conference on Learning Representations , 2023
2023
Closest in time.
Y. Yue, R. Lu, B. Kang, S. Song, and G. Huang, “Understanding, predicting and better resolving q-value divergence in offline-RL,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Closest in time.
L. Zhang, Y. Feng, R. Wang, Y. Xu, N. Xu, Z. Liu, and H. Du, “Efficient experience replay architecture for offline reinforcement learning,” Robotic Intelligence and Automation , 2023
2023
Closest in time.
H. Chen, C. Lu, C. Ying, H. Su, and J. Zhu, “Offline reinforcement learning via high-fidelity generative behavior modeling,” in The Eleventh International Conference on Learning Representations , 2023
2023
Closest in time.
Q. Yang, S. Wang, Q. Zhang, G. Huang, and S. Song, “Hundreds guide millions: Adaptive offline reinforcement learning with expert guidance,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Closest in time.
J. Fu, A. Kumar, M. Soh, and S. Levine, “Diagnosing bottlenecks in deep q-learning algorithms,” in International Conference on Machine Learning , 2019, pp. 2021–2030
2030
Closest in time.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in International conference on machine learning , 2019, pp. 2052–2062
2062
Closest in time.