Fetching the paper…
Reading the bibliography…
In this work, we argue for the importance of an online evaluation budget for a reliable comparison of deep offline RL algorithms.
Efficient Training of Artificial Neural Networks for Autonomous Navigation
Pomerleau, D. A · 1991
Earlier work this paper cites.
D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2004
Earlier work this paper cites.
A benchmark environment motivated by industrial control problems
Hein, D., Depeweg, S., Tokic, M., Udluft, S., Hentschel, A., Runkler, T. A., and Sterzing, V · 2017
Earlier work this paper cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Deep Reinforcement Learning That Matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Earlier work this paper cites.
Top-K Off-Policy Correction for a REINFORCE Recommender System
Chen, M., Beutel, A., Covington, P., Jain, S., Belletti, F., and Chi, E. H · 2019
Earlier work this paper cites.
Show Your Work: Improved Reporting of Experimental Results
Dodge, J., Gururangan, S., Card, D., Schwartz, R., and Smith, N. A · 2019
Earlier work this paper cites.
Benchmarking Batch Deep Reinforcement Learning Algorithms
Fujimoto, S., Conti, E., Ghavamzadeh, M., and Pineau, J · 2019
Earlier work this paper cites.
Horizon: Facebook’s Open Source Applied Reinforcement Learning Platform
Gauci, J., Conti, E., Liang, Y., Virochsiri, K., He, Y., Kaden, Z., Narayanan, V., Ye, X., Chen, Z., and Fujimoto, S · 2019
Earlier work this paper cites.
Off-Policy Evaluation via Off-Policy Classification
Irpan, A., Rao, K., Bousmalis, K., Harris, C., Ibarz, J., and Levine, S · 2019
Earlier work this paper cites.
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
Kumar, A., Fu, J., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Batch Policy Learning under Constraints
Le, H. M., Voloshin, C., and Yue, Y · 2019
Earlier work this paper cites.
A Deep Value-network Based Approach for Multi-Driver Order Dispatching
Tang, X., Qin, Z. T., Zhang, F., Wang, Z., Xu, Z., Ma, Y., Zhu, H., and Ye, J · 2019
Earlier work this paper cites.
CityLearn v1.0: An OpenAI Gym Environment for Demand Response with Deep Reinforcement Learning
Vázquez-Canteli, J. R., Kämpf, J., Henze, G., and Nagy, Z · 2019
Cited alongside, same era.
Behavior Regularized Offline Reinforcement Learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
Conservative Q-Learning for Offline Reinforcement Learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
FinRL: A Deep Reinforcement Learning Library for Automated Stock Trading in Quantitative Finance
MOPO: Model-based Offline Policy Optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.
PLAS: Latent Action Space for Offline Reinforcement Learning
Zhou, W., Bajracharya, S., and Held, D · 2020
Later among the works it cites.
A Minimalist Approach to Offline Reinforcement Learning
Fujimoto, S. and Gu, S. S · 2021
Closest in time.
RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning
Gulcehre, C., Wang, Z., Novikov, A., Paine, T. L., Colmenarejo, S. G., Zolna, K., Agarwal, R., Merel, J., Mankowitz, D., Paduraru, C., Dulac-Arnold, G., Li, J., Norouzi, M., Hoffman, M., Nachum, O., Tucker, G., Heess, N., and de Freitas, N · 2021
Closest in time.
Hyperparameter Selection for Imitation Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, X.-Y., Yang, H., Chen, Q., Zhang, R., Yang, L., Xiao, B., and Wang, C. D · 2020
Cited alongside, same era.
Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
Matsushima, T., Furuta, H., Matsuo, Y., Nachum, O., and Gu, S · 2020
Cited alongside, same era.
Hyperparameter Selection for Offline Reinforcement Learning
Paine, T. L., Paduraru, C., Michi, A., Gulcehre, C., Zolna, K., Novikov, A., Wang, Z., and de Freitas, N · 2020
Cited alongside, same era.
Testing machine learning based systems: a systematic mapping
Riccio, V., Jahangirova, G., Stocco, A., Humbatova, N., Weiss, M., and Tonella, P · 2020
Cited alongside, same era.
Showing Your Work Doesn’t Always Work
Tang, R. et al · 2020
Cited alongside, same era.
Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning
Voloshin, C., Le, H. M., Jiang, N., and Yue, Y · 2020
Cited alongside, same era.
Wang, Z., Novikov, A., Zolna, K., Springenberg, J. T., Reed, S., Shahriari, B., Siegel, N., Merel, J., Gulcehre, C., Heess, N., and de Freitas, N · 2020
Cited alongside, same era.
Self-Supervised Reinforcement Learning for Recommender Systems
Xin, X., Karatzoglou, A., Arapakis, I., and Jose, J. M · 2020
Cited alongside, same era.
Hussenot, L., Andrychowicz, M., Vincent, D., Dadashi, R., Raichuk, A., Ramos, S., Momchev, N., Girgin, S., Marinier, R., Stafiniak, L., Orsini, M., Bachem, O., Geist, M., and Pietquin, O · 2021
Closest in time.
Active Offline Policy Selection
Konyushkova, K., Chen, Y., Paine, T. L., Gulcehre, C., Paduraru, C., Mankowitz, D. J., Denil, M., and de Freitas, N · 2021
Closest in time.
Offline Reinforcement Learning with Fisher Divergence Critic Regularization
Kostrikov, I., Tompson, J., Fergus, R., and Nachum, O · 2021
Closest in time.
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Martín-Martín, R · 2021
Closest in time.
NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning
Qin, R., Gao, S., Zhang, X., Xu, Z., Huang, S., Li, Z., Zhang, W., and Yu, Y · 2021
Closest in time.
Su, D., Lee, J. D., Mulvey, J. M., and Poor, H. V · 2021
Closest in time.
Deep Reinforcement Learning at the Edge of the Statistical Precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2022
Closest in time.
Expected Validation Performance and Estimation of a Random Variable’s Maximum
Dodge, J. et al · 2022
Closest in time.