Fetching the paper…
Reading the bibliography…
Off-policy reinforcement learning (RL) has achieved notable success in tackling many complex real-world tasks, by leveraging previously collected data for policy learning.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Contraction mappings in the theory underlying dynamic programming
Denardo, E. V · 1967
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Active lmitation learning: formal and practical reductions to iid learning
Judah, K., Fern, A. P., Dietterich, T. G., and Tadepalli, P · 2014
Earlier work this paper cites.
Mujoco haptix: A virtual reality system for hand manipulation
Kumar, V. and Todorov, E · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D · 2019
Earlier work this paper cites.
Reinforcement learning with combinatorial actions: An application to vehicle routing
Delarue, A., Anderson, R., and Tjandraatmadja, C · 2020
Earlier work this paper cites.
Sharing knowledge in multi-task deep reinforcement learning
DEramo, C., Tateo, D., Bonarini, A., Restelli, M., and Peters, J · 2020
Earlier work this paper cites.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y., Jia, Z., and Wang, M · 2020
Cited alongside, same era.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N. and Uehara, M · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Lee, A. X., Nagabandi, A., Abbeel, P., and Levine, S · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Understanding deep neural function approximation in reinforcement learning via ϵ \epsilon -greedy exploration
Liu, F., Viano, L., and Cevher, V · 2022
Later among the works it cites.
When to trust your simulator: Dynamics-aware hybrid offline-and-online reinforcement learning
Niu, H., Qiu, Y., Li, M., Zhou, G., HU, J., Zhan, X., et al · 2022
Later among the works it cites.
Robust reinforcement learning using offline data
Panaganti, K., Xu, Z., Kalathil, D., and Ghavamzadeh, M · 2022
Later among the works it cites.
Mastering the game of stratego with model-free multiagent reinforcement learning
Perolat, J., De Vylder, B., Hennes, D., Tarassov, E., Strub, F., de Boer, V., Muller, P., Connor, J. T., Burch, N., Anthony, T., et al · 2022
Later among the works it cites.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Cited alongside, same era.
Deep reinforcement learning for autonomous driving: A survey
Kiran, B. R., Sobh, I., Talpaert, V., Mannion, P., Al Sallab, A. A., Yogamani, S., and Pérez, P · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Cited alongside, same era.
A graph placement methodology for fast chip design
Mirhoseini, A., Goldie, A., Yazgan, M., Jiang, J. W., Songhori, E., Wang, S., Lee, Y.-J., Johnson, E., Pathak, O., Nazi, A., et al · 2021
Cited alongside, same era.
Tactical optimism and pessimism for deep reinforcement learning
Moskovitz, T., Parker-Holder, J., Pacchiano, A., Arbel, M., and Jordan, M · 2021
Cited alongside, same era.
Hybrid rl: Using both offline and online data can make rl efficient
Song, Y., Zhou, Y., Sekhari, A., Bagnell, D., Krishnamurthy, A., and Sun, W · 2022
Later among the works it cites.
Deep reinforcement learning: a survey
Wang, X., Wang, S., Liang, X., Zhao, D., Huang, J., Xu, X., Dai, B., and Miao, Q · 2022
Later among the works it cites.
Controlling underestimation bias in reinforcement learning via quasi-median operation
Wei, W., Zhang, Y., Liang, J., Li, L., and Li, Y · 2022
Later among the works it cites.
The in-sample softmax for offline reinforcement learning
Xiao, C., Wang, H., Pan, Y., White, A., and White, M · 2022
Later among the works it cites.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S · 2023
Later among the works it cites.
Seizing serendipity: Exploiting the value of past success in off-policy actor-critic
Ji, T., Luo, Y., Sun, F., Zhan, X., Zhang, J., and Xu, H · 2023
Later among the works it cites.
Model-based reinforcement learning: A survey
Moerland, T. M., Broekens, J., Plaat, A., Jonker, C. M., et al · 2023
Later among the works it cites.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Prudencio, R. F., Maximo, M. R., and Colombini, E. L · 2023
Later among the works it cites.
Jump-start reinforcement learning
Uchendu, I., Xiao, T., Lu, Y., Zhu, B., Yan, M., Simon, J., Bennice, M., Fu, C., Ma, C., Jiao, J., et al · 2023
Later among the works it cites.
Leveraging offline data in online reinforcement learning
Wagenmaker, A. and Pacchiano, A · 2023
Later among the works it cites.
Actor-critic alignment for offline-to-online reinforcement learning
Yu, Z. and Zhang, X · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control
Hansen, N., Su, H., and Wang, X · 2024
Closest in time.
Odice: Revealing the mystery of distribution correction estimation via orthogonal-gradient update
Mao, L., Xu, H., Zhang, W., and Zhan, X · 2024
Closest in time.
Drm: Mastering visual reinforcement learning through dormant ratio minimization
Xu, G., Zheng, R., Liang, Y., Wang, X., Yuan, Z., Ji, T., Luo, Y., Liu, X., Yuan, J., Hua, P., et al · 2024
Closest in time.