Fetching the paper…
Reading the bibliography…
Deep reinforcement learning for continuous control has recently achieved impressive progress.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Maxmin q-learning: Controlling the estimation bias of q-learning, 2021
Lan, Q., Pan, Y., Fyshe, A., and White, M · 2002
Earlier work this paper cites.
Reinforcement learning with combinatorial actions: An application to vehicle routing, 2020
Delarue, A., Anderson, R., and Tjandraatmadja, C · 2010
Earlier work this paper cites.
On the role of planning in model-based deep reinforcement learning, 2021
Hamrick, J. B., Friesen, A. L., Behbahani, F., Guez, A., Viola, F., Witherspoon, S., Anthony, T., Buesing, L., Veličković, P., and Weber, T · 2011
Earlier work this paper cites.
Infantile amnesia: a neurogenic hypothesis
Josselyn, S. A. and Frankland, P. W · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Hippocampal neurogenesis regulates forgetting during adulthood and infancy
Akers, K. G., Martinez-Canabal, A., Restivo, L., Yiu, A. P., De Cristofaro, A., Hsiang, H.-L., Wheeler, A. L., Guskjolen, A., Niibori, Y., Shoji, H., et al · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Earlier work this paper cites.
Prioritized experience replay, 2016
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Infantile amnesia: a critical period of learning to learn and remember
Alberini, C. M. and Travaglia, A · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically, 2017
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Hindsight experience replay, 2018
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Cited alongside, same era.
Deep reinforcement learning and the deadly triad, 2018
van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Cited alongside, same era.
A deeper look at experience replay, 2018
Zhang, S. and Sutton, R. S · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration, 2019
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Revisiting fundamentals of experience replay
Fedus, W., Ramachandran, P., Agarwal, R., Bengio, Y., Larochelle, H., Rowland, M., and Dabney, W · 2020
Deep reinforcement learning with plasticity injection, 2023
Nikishin, E., Oh, J., Ostrovski, G., Lyle, C., Pascanu, R., Dabney, W., and Barreto, A · 2023
Later among the works it cites.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Prudencio, R. F., Maximo, M. R., and Colombini, E. L · 2023
Later among the works it cites.
The primacy bias in model-based rl
Qiao, Z., Lyu, J., and Li, X · 2023
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency, 2023
Schwarzer, M., Obando-Ceron, J., Courville, A., Bellemare, M., Agarwal, R., and Castro, P. S · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S., and Evci, U · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D2rl: Deep dense architectures in reinforcement learning
Sinha, S., Bharadhwaj, H., Srinivas, A., and Garg, A · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Randomized ensembled double q-learning: Learning fast without a model, 2021
Chen, X., Wang, C., Zhou, Z., and Ross, K · 2021
Cited alongside, same era.
Data-driven human-robot interaction without velocity measurement using off-policy reinforcement learning
Yang, Y., Ding, Z., Wang, R., Modares, H., and Wunsch, D. C · 2021
Cited alongside, same era.
Towards deeper deep reinforcement learning with spectral normalization, 2022
Bjorck, J., Gomes, C. P., and Weinberger, K. Q · 2022
Cited alongside, same era.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A · 2022
Cited alongside, same era.
Multi-game decision transformers, 2022
Lee, K.-H., Nachum, O., Yang, M., Lee, L., Freeman, D., Xu, W., Guadarrama, S., Fischer, I., Jang, E., Michalewski, H., and Mordatch, I · 2022
Cited alongside, same era.
Later among the works it cites.
Drm: Mastering visual reinforcement learning through dormant ratio minimization
Xu, G., Zheng, R., Liang, Y., Wang, X., Yuan, Z., Ji, T., Luo, Y., Liu, X., Yuan, J., Hua, P., et al · 2023
Later among the works it cites.
Mastering diverse domains through world models, 2024
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2024
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control, 2024
Hansen, N., Su, H., and Wang, X · 2024
Later among the works it cites.
Simba: Simplicity bias for scaling up parameters in deep reinforcement learning
Lee, H., Hwang, D., Kim, D., Kim, H., Tai, J. J., Subramanian, K., Wurman, P. R., Choo, J., Stone, P., and Seno, T · 2024
Later among the works it cites.
Neuroplastic expansion in deep reinforcement learning
Liu, J., Obando-Ceron, J., Courville, A., and Pan, L · 2024
Later among the works it cites.
Offline-boosted actor-critic: Adaptively blending optimal historical behaviors in deep off-policy rl
Luo, Y., Ji, T., Sun, F., Zhang, J., Xu, H., and Zhan, X · 2024
Later among the works it cites.
Learning better with less: effective augmentation for sample-efficient visual reinforcement learning
Ma, G., Zhang, L., Wang, H., Li, L., Wang, Z., Wang, Z., Shen, L., Wang, X., and Tao, D · 2024
Later among the works it cites.
Humanoidbench: Simulated humanoid benchmark for whole-body locomotion and manipulation, 2024
Sferrazza, C., Huang, D.-M., Lin, X., Lee, Y., and Abbeel, P · 2024
Later among the works it cites.
Model-based off-policy deep reinforcement learning with model-embedding
Tan, X., Qu, C., Xiong, J., Zhang, J., Qiu, X., and Jin, Y · 2024
Later among the works it cites.
Yenicesu, A. S., Mutlu, F. B., Kozat, S. S., and Oguz, O. S · 2024
Later among the works it cites.