Fetching the paper…
Reading the bibliography…
Reinforcement learning has achieved significant milestones, but sample efficiency remains a bottleneck for real-world applications.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
T. Van Laarhoven · 2017
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare · 2021
Cited alongside, same era.
Randomized ensembled double Q-learning: Learning fast without a model
X. Chen, C. Wang, Z. Zhou, and K. Ross · 2021
Cited alongside, same era.
Dropout q-functions for doubly efficient reinforcement learning
T. Hiraoka, T. Imagawa, T. Hashimoto, T. Onishi, and Y. Tsuruoka · 2021
Cited alongside, same era.
Samba: Safe model-based & active reinforcement learning
A. I. Cowen-Rivers, D. Palenicek, V. Moens, M. A. Abdullah, A. Sootla, J. Wang, and H. Bou-Ammar · 2022
Cited alongside, same era.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
P. D’Oro, M. Schwarzer, E. Nikishin, P.-L. Bacon, M. G. Bellemare, and A. Courville · 2022
Cited alongside, same era.
The primacy bias in deep reinforcement learning
Normalization techniques in training dnns: Methodology, analysis and application
L. Huang, J. Qin, Y. Zhou, F. Zhu, L. Liu, and L. Shao · 2023
Later among the works it cites.
Efficient deep reinforcement learning requires regulating overfitting
Q. Li, A. Kumar, I. Kostrikov, and S. Levine · 2023
Later among the works it cites.
CrossQ: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
A. Bhatt, D. Palenicek, B. Belousov, M. Argus, A. Amiranashvili, T. Brox, and J. Peters · 2024
Later among the works it cites.
Dissecting deep rl with high update ratios: Combatting value overestimation and divergence
M. Hussing, C. Voelcker, I. Gilitschenski, A.-m. Farahmand, and E. Eaton · 2024
Later among the works it cites.
Normalization and effective learning rates in reinforcement learning
C. Lyle, Z. Zheng, K. Khetarpal, J. Martens, H. van Hasselt, R. Pascanu, and W. Dabney · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Nikishin, M. Schwarzer, P. D’Oro, P.-L. Bacon, and A. Courville · 2022
Cited alongside, same era.
Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control
M. Nauman, M. Ostaszewski, K. Jankowski, P. Miłoś, and M. Cygan · 2024
Later among the works it cites.