Fetching the paper…
Reading the bibliography…
To obtain a near-optimal policy with fewer interactions in Reinforcement Learning (RL), a promising approach involves the combination of offline RL, which enhances sample efficiency by leveraging offline datasets, and online RL, which explores informative transitions by interacting with the environment.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., & Nachum, O. (2019) · 1911
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., & Zhang, S. (2019) · 1912
Earlier work this paper cites.
Maxmin Q-learning: Controlling the estimation bias of Q-learning
Lan, Q., Pan, Y., Fyshe, A., & White, M. (2020) · 2002
Earlier work this paper cites.
D4RL: datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., & Levine, S. (2020) · 2004
Earlier work this paper cites.
AWAC: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., & Levine, S. (2020) · 2006
Earlier work this paper cites.
Uncertainty propagation for quality assurance in reinforcement learning
Schneegass, D., Udluft, S., & Martinetz, T. (2008) · 2008
Earlier work this paper cites.
Visualizing data using t-SNE.
Van der Maaten, L., & Hinton, G. (2008) · 2008
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., & Riedmiller, M. (2012) · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P. (2015) · 2015
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., & Van Roy, B. (2016) · 2016
Earlier work this paper cites.
Ucb exploration via Q-ensembles
Chen, R. Y., Sidor, S., Abbeel, P., & Schulman, J. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017) · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., & Abbeel, P. (2017) · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., & Meger, D. (2018) · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018) · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., & Silver, D. (2018) · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2018) · 2018
Cited alongside, same era.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L., & Wang, M. (2019) · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., & Jordan, M. I. (2020) · 2020
Cited alongside, same era.
Conservative Q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020) · 2020
Cited alongside, same era.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Bai, C., Wang, L., Yang, Z., Deng, Z.-H., Garg, A., Liu, P., & Wang, Z. (2022) · 2022
Later among the works it cites.
Why so pessimistic? estimating uncertainties for offline RL through ensembles, and why their independence matters
Ghasemipour, K., Gu, S. S., & Nachum, O. (2022) · 2022
Later among the works it cites.
Offline reinforcement learning with implicit Q-learning
Kostrikov, I., Nair, A., & Levine, S. (2022) · 2022
Later among the works it cites.
The challenges of exploration for offline reinforcement learning
Lambert, N., Wulfmeier, M., Whitney, W., Byravan, A., Bloesch, M., Dasagi, V., Hertweck, T., & Riedmiller, M. (2022) · 2022
Later among the works it cites.
Offline-to-online reinforcement learning via balanced replay and pessimistic Q-ensemble
Lee, S., Seo, Y., Lee, K., Abbeel, P., & Shin, J. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shen, Q., Li, Y., Jiang, H., Wang, Z., & Zhao, T. (2020) · 2020
Cited alongside, same era.
MOPO: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., & Ma, T. (2020) · 2020
Cited alongside, same era.
Uncertainty-based offline reinforcement learning with diversified Q-ensemble
An, G., Moon, S., Kim, J.-H., & Song, H. O. (2021) · 2021
Cited alongside, same era.
Randomized ensembled double Q-learning: Learning fast without a model
Chen, X., Wang, C., Zhou, Z., & Ross, K. W. (2021) · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Fujimoto, S., & Gu, S. S. (2021) · 2021
Cited alongside, same era.
Is pessimism provably efficient for offline RL?
Jin, Y., Yang, Z., & Wang, Z. (2021) · 2021
Cited alongside, same era.
Deep reinforcement learning for autonomous driving: A survey
Kiran, B. R., Sobh, I., Talpaert, V., Mannion, P., Al Sallab, A. A., Yogamani, S., & Pérez, P. (2021) · 2021
Cited alongside, same era.
A dataset perspective on offline reinforcement learning
Schweighofer, K., Dinu, M.-c., Radler, A., Hofmarcher, M., Patil, V. P., Bitto-Nemling, A., Eghbal-zadeh, H., & Hochreiter, S. (2022) · 2022
Later among the works it cites.
S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics
Sinha, S., Mandlekar, A., & Garg, A. (2022) · 2022
Later among the works it cites.
Supported policy optimization for offline reinforcement learning
Wu, J., Wu, H., Qiu, Z., Wang, J., & Long, M. (2022) · 2022
Later among the works it cites.
RORL: Robust offline reinforcement learning via conservative smoothing
Yang, R., Bai, C., Ma, X., Wang, Z., Zhang, C., & Han, L. (2022) · 2022
Later among the works it cites.
Adaptive behavior cloning regularization for stable offline-to-online reinforcement learning
Zhao, Y., Boney, R., Ilin, A., Kannala, J., & Pajarinen, J. (2022) · 2022
Later among the works it cites.
Cal-QL: Calibrated offline RL pre-training for efficient online fine-tuning
Nakamoto, M., Zhai, Y., Singh, A., Mark, M. S., Ma, Y., Finn, C., Kumar, A., & Levine, S. (2023) · 2023
Later among the works it cites.
Jump-start reinforcement learning
Uchendu, I., Xiao, T., Lu, Y., Zhu, B., Yan, M., Simon, J., Bennice, M., Fu, C., Ma, C., Jiao, J., et al. (2023) · 2023
Later among the works it cites.
Policy expansion for bridging offline-to-online reinforcement learning
Zhang, H., Xu, W., & Yu, H. (2023) · 2023
Later among the works it cites.
Improving offline-to-online reinforcement learning with Q-ensembles
Zhao, K., Ma, Y., Liu, J., Jianye, H., Zheng, Y., & Meng, Z. (2023) · 2023
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., & Precup, D. (2019) · 2062
Later among the works it cites.