Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) harnesses the power of massive datasets for resolving sequential decision problems.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic Policy Gradient Algorithms,” in Proceedings of the 31st International Conference on Machine Learning , Jan. 2014, pp. 387–395
2014
Earlier work this paper cites.
B. Scherrer, M. Ghavamzadeh, V. Gabillon, B. Lesner, and M. Geist, “Approximate Modified Policy Iteration and its Application to the Game of Tetris,” Journal of Machine Learning Research , vol. 16, no. 49, pp. 1629–1676, 2015
2015
Earlier work this paper cites.
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy, “Deep Exploration via Bootstrapped DQN,” in Advances in Neural Information Processing Systems , vol. 29, 2016
2016
Earlier work this paper cites.
B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,” in Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
I. Osband, J. Aslanides, and A. Cassirer, “Randomized Prior Functions for Deep Reinforcement Learning,” in Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
M. Farajtabar, Y. Chow, and M. Ghavamzadeh, “More Robust Doubly Robust Off-policy Evaluation,” in Proceedings of the 35th International Conference on Machine Learning , Jul. 2018, pp. 1447–1456
2018
Earlier work this paper cites.
Q. Liu, L. Li, Z. Tang, and D. Zhou, “Breaking the Curse of Horizon: Infinite-Horizon Off-Policy Estimation,” in Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
S. Fujimoto, H. Hoof, and D. Meger, “Addressing Function Approximation Error in Actor-Critic Methods,” in Proceedings of the 35th International Conference on Machine Learning , Jul. 2018, pp. 1587–1596
2018
Earlier work this paper cites.
Y. Wu, G. Tucker, and O. Nachum, “Behavior Regularized Offline Reinforcement Learning,” Nov. 2019
2019
Earlier work this paper cites.
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine, “Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction,” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
K. Ciosek, Q. Vuong, R. Loftin, and K. Hofmann, “Better Exploration with Optimistic Actor Critic,” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
S. Fujimoto, E. Conti, M. Ghavamzadeh, and J. Pineau, “Benchmarking Batch Deep Reinforcement Learning Algorithms,” Oct. 2019
2019
Earlier work this paper cites.
T. Xie, Y. Ma, and Y.-X. Wang, “Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling,” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
O. Nachum, Y. Chow, B. Dai, and L. Li, “DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections,” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Chen and N. Jiang, “Information-Theoretic Considerations in Batch Reinforcement Learning,” in Proceedings of the 36th International Conference on Machine Learning , May 2019, pp. 1042–1051
2019
Earlier work this paper cites.
S. Levine, A. Kumar, G. Tucker, and J. Fu, “Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems,” Nov. 2020
2020
Cited alongside, same era.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative Q-Learning for Offline Reinforcement Learning,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 1179–1191
2020
Cited alongside, same era.
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma, “MOPO: Model-based Offline Policy Optimization,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 14 129–14 142
2020
Cited alongside, same era.
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims, “MOReL: Model-Based Offline Reinforcement Learning,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 21 810–21 823
2020
Cited alongside, same era.
S. K. S. Ghasemipour, D. Schuurmans, and S. S. Gu, “EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RL,” in Proceedings of the 38th International Conference on Machine Learning , Jul. 2021, pp. 3682–3691
2021
Closest in time.
A. Nair, A. Gupta, M. Dalal, and S. Levine, “AWAC: Accelerating Online Reinforcement Learning with Offline Datasets,” Apr. 2021
2021
Closest in time.
I. Kostrikov, R. Fergus, J. Tompson, and O. Nachum, “Offline Reinforcement Learning with Fisher Divergence Critic Regularization,” in Proceedings of the 38th International Conference on Machine Learning , Jul. 2021, pp. 5774–5783
2021
Closest in time.
Y. Wu, S. Zhai, N. Srivastava, J. M. Susskind, J. Zhang, R. Salakhutdinov, and H. Goh, “Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning,” in Proceedings of the 38th International Conference on Machine Learning , Jul. 2021, pp. 11 319–11 328
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Jiang and J. Huang, “Minimax Value Interval for Off-Policy Evaluation and Policy Optimization,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 2747–2758
2020
Cited alongside, same era.
Y. Duan, Z. Jia, and M. Wang, “Minimax-Optimal Off-Policy Evaluation with Linear Function Approximation,” in Proceedings of the 37th International Conference on Machine Learning , Nov. 2020, pp. 2701–2709
2020
Cited alongside, same era.
M. Yang, O. Nachum, B. Dai, L. Li, and D. Schuurmans, “Off-Policy Evaluation via the Regularized Lagrangian,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 6551–6561
2020
Cited alongside, same era.
T. Xie and N. Jiang, “Q* Approximation Schemes for Batch Reinforcement Learning: A Theoretical Comparison,” in Proceedings of the 36th Conference on Uncertainty in Artificial Intelligence (UAI) , Aug. 2020, pp. 550–559
2020
Cited alongside, same era.
J. Fan, Z. Wang, Y. Xie, and Z. Yang, “A Theoretical Analysis of Deep Q-Learning,” in Proceedings of the 2nd Conference on Learning for Dynamics and Control , Jul. 2020, pp. 486–489
2020
Cited alongside, same era.
Q. Cai, Z. Yang, C. Jin, and Z. Wang, “Provably Efficient Exploration in Policy Optimization,” in Proceedings of the 37th International Conference on Machine Learning , Nov. 2020, pp. 1283–1294
2020
Cited alongside, same era.
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine, “D4RL: Datasets for Deep Data-Driven Reinforcement Learning,” Feb. 2021
2021
Cited alongside, same era.
Y. Jin, Z. Yang, and Z. Wang, “Is Pessimism Provably Efficient for Offline RL?” in Proceedings of the 38th International Conference on Machine Learning , Jul. 2021, pp. 5084–5096
2021
Cited alongside, same era.
S. Fujimoto and S. S. Gu, “A Minimalist Approach to Offline Reinforcement Learning,” in Advances in Neural Information Processing Systems , vol. 34, 2021, pp. 20 132–20 145
2021
Closest in time.
G. An, S. Moon, J.-H. Kim, and H. O. Song, “Uncertainty-Based Offline Reinforcement Learning with Diversified Q-Ensemble,” in Advances in Neural Information Processing Systems , vol. 34, 2021, pp. 7436–7447
2021
Closest in time.
M. Yin, Y. Bai, and Y.-X. Wang, “Near-Optimal Provable Uniform Convergence in Offline Policy Evaluation for Reinforcement Learning,” in Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , Mar. 2021, pp. 1567–1575
2021
Closest in time.
T. Xie and N. Jiang, “Batch Value-function Approximation with Only Realizability,” in Proceedings of the 38th International Conference on Machine Learning . PMLR, Jul. 2021, pp. 11 404–11 413
2021
Closest in time.
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare, “Deep Reinforcement Learning at the Edge of the Statistical Precipice,” in Advances in Neural Information Processing Systems , vol. 34, 2021, pp. 29 304–29 320
2021
Closest in time.
C. Bai, L. Wang, Z. Yang, Z.-H. Deng, A. Garg, P. Liu, and Z. Wang, “Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning,” in International Conference on Learning Representations , Mar. 2022
2022
Closest in time.
Y. Guo, S. Feng, N. L. Roux, E. Chi, H. Lee, and M. Chen, “Batch Reinforcement Learning Through Continuation Method,” in International Conference on Learning Representations , Feb. 2022
2022
Closest in time.
R. Zhang, B. Dai, L. Li, and D. Schuurmans, “GenDICE: Generalized Offline Estimation of Stationary Values,” in International Conference on Learning Representations , Feb. 2022
2022
Closest in time.
P. Liao, Z. Qi, R. Wan, P. Klasnja, and S. Murphy, “Batch Policy Learning in Average Reward Markov Decision Processes,” Sep. 2022
2022
Closest in time.
S. Fujimoto, D. Meger, and D. Precup, “Off-Policy Deep Reinforcement Learning without Exploration,” in Proceedings of the 36th International Conference on Machine Learning , May 2019, pp. 2052–2062
2062
Closest in time.