Fetching the paper…
Reading the bibliography…
Recent advancements in off-policy Reinforcement Learning (RL) have significantly improved sample efficiency, primarily due to the incorporation of various forms of regularization that enable more gradient update steps than traditional agents.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Critical learning periods in deep neural networks
Achille, A., Rovere, M., and Soatto, S · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
Self-attention generative adversarial networks
Zhang, H., Goodfellow, I., Metaxas, D., and Odena, A · 2018
Earlier work this paper cites.
Better exploration with optimistic actor critic
Ciosek, K., Vuong, Q., Loftin, R., and Hofmann, K · 2019
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
OpenAI, :, Berner, C., et al · 2019
Cited alongside, same era.
The bitter lesson
Sutton, R · 2019
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Cited alongside, same era.
On warm-starting neural network training
Ash, J. and Adams, R. P · 2020
Cited alongside, same era.
Randomized ensembled double q-learning: Learning fast without a model
Chen, X., Wang, C., Zhou, Z., and Ross, K. W · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2020
Cited alongside, same era.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A · 2022
Later among the works it cites.
Temporal difference learning for model predictive control
Hansen, N., Wang, X., and Su, H · 2022
Later among the works it cites.
Efficient deep reinforcement learning requires regulating overfitting
Li, Q., Kumar, A., Kostrikov, I., and Levine, S · 2022
Later among the works it cites.
Understanding and preventing capacity loss in reinforcement learning
Lyle, C., Rowland, M., and Dabney, W · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L., and Courville, A · 2022
Later among the works it cites.
A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Kumar, A., Agarwal, R., Ghosh, D., and Levine, S · 2020
Cited alongside, same era.
Regularization matters in policy optimization-an empirical study on continuous control
Liu, Z., Li, X., Kang, B., and Darrell, T · 2020
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2021
Cited alongside, same era.
What matters in on-policy reinforcement learning? a large-scale empirical study
Andrychowicz, M., Raichuk, A., Stańczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., et al · 2021
Cited alongside, same era.
Towards deeper deep reinforcement learning with spectral normalization
Bjorck, J., Gomes, C. P., and Weinberger, K. Q · 2021
Cited alongside, same era.
Smith, L., Kostrikov, I., and Levine, S · 2022
Later among the works it cites.
Loss of plasticity in continual deep reinforcement learning
Abbas, Z., Zhao, R., Modayil, J., White, A., and Machado, M. C · 2023
Later among the works it cites.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S · 2023
Later among the works it cites.
Learning pessimism for reinforcement learning
Cetin, E. and Celiktutan, O · 2023
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Later among the works it cites.
Maintaining plasticity in continual learning via regenerative regularization
Kumar, S., Marklund, H., and Roy, B. V · 2023
Later among the works it cites.
Plastic: Improving input and label plasticity for sample efficient reinforcement learning
Lee, H., Cho, H., Kim, H., Gwak, D., Kim, J., Choo, J., Yun, S.-Y., and Yun, C · 2023
Later among the works it cites.
Understanding plasticity in neural networks
Lyle, C., Zheng, Z., Nikishin, E., Avila Pires, B., Pascanu, R., and Dabney, W · 2023
Later among the works it cites.
Deep reinforcement learning with plasticity injection
Nikishin, E., Oh, J., Ostrovski, G., Lyle, C., Pascanu, R., Dabney, W., and Barreto, A · 2023
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency
Schwarzer, M., Ceron, J. S. O., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S., and Evci, U · 2023
Later among the works it cites.
Dissecting deep rl with high update ratios: Combatting value overestimation and divergence
Hussing, M., Voelcker, C., Gilitschenski, I., Farahmand, A.-m., and Eaton, E · 2024
Closest in time.
Seizing serendipity: Exploiting the value of past success in off-policy actor-critic, 2024
Ji, T., Luo, Y., Sun, F., Zhan, X., Zhang, J., and Xu, H · 2024
Closest in time.