Fetching the paper…
Reading the bibliography…
Reinforcement learning has achieved significant milestones, but sample efficiency remains a bottleneck for real-world applications.
On warm-starting neural network training, 2020
Jordan T. Ash and Ryan P. Adams · 1910
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Earlier work this paper cites.
Wenling Shang, Kihyuk Sohn, Diogo Almeida, and Honglak Lee · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I Loshchilov · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
Twan Van Laarhoven · 2017
Earlier work this paper cites.
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I. Jordan, Joseph E. Gonzalez, and Sergey Levine · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research, 2018
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, Vikash Kumar, and Wojciech Zaremba · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Deepmellow: Removing the need for a target network in deep q-learning
Seungchan Kim, Kavosh Asadi, Michael Littman, and George Konidaris · 2019
Cited alongside, same era.
When to use parametric models in reinforcement learning?
Hado P Van Hasselt, Matteo Hessel, and John Aslanides · 2019
Cited alongside, same era.
Root mean square layer normalization
Sample-efficient reinforcement learning by breaking the replay ratio barrier
Pierluca D’Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon, Marc G Bellemare, and Aaron Courville · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Later among the works it cites.
Normalization techniques in training dnns: Methodology, analysis and application
Lei Huang, Jie Qin, Yi Zhou, Fan Zhu, Li Liu, and Ling Shao · 2023
Later among the works it cites.
Plastic: Improving input and label plasticity for sample efficient reinforcement learning, 2023
Hojoon Lee, Hanseul Cho, Hyunseung Kim, Daehoon Gwak, Joonkee Kim, Jaegul Choo, Se-Young Yun, and Chulhee Yun · 2023
Later among the works it cites.
Efficient deep reinforcement learning requires regulating overfitting
Qiyang Li, Aviral Kumar, Ilya Kostrikov, and Sergey Levine · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Biao Zhang and Rico Sennrich · 2019
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Cited alongside, same era.
Grokking deep reinforcement learning
Miguel Morales · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Cited alongside, same era.
Randomized ensembled double Q-learning: Learning fast without a model
Xinyue Chen, Che Wang, Zijian Zhou, and Keith Ross · 2021
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Cited alongside, same era.
Dropout q-functions for doubly efficient reinforcement learning
Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, and Yoshimasa Tsuruoka · 2021
Cited alongside, same era.
Bigger, better, faster: Human-level atari with human-level efficiency, 2023
Max Schwarzer, Johan Obando-Ceron, Aaron Courville, Marc Bellemare, Rishabh Agarwal, and Pablo Samuel Castro · 2023
Later among the works it cites.
CrossQ: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
Aditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus, Artemij Amiranashvili, Thomas Brox, and Jan Peters · 2024
Later among the works it cites.
Weight clipping for deep continual and reinforcement learning
Mohamed Elsayed, Qingfeng Lan, Clare Lyle, and A Rupam Mahmood · 2024
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control, 2024
Nicklas Hansen, Hao Su, and Xiaolong Wang · 2024
Later among the works it cites.
Dissecting deep rl with high update ratios: Combatting value overestimation and divergence
Marcel Hussing, Claas Voelcker, Igor Gilitschenski, Amir-massoud Farahmand, and Eric Eaton · 2024
Later among the works it cites.
Normalization and effective learning rates in reinforcement learning
Clare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens, Hado van Hasselt, Razvan Pascanu, and s Will Dabney · 2024
Later among the works it cites.
Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control
Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Miłoś, and Marek Cygan · 2024
Later among the works it cites.
Gait in eight: Efficient on-robot learning for omnidirectional quadruped locomotion
Nico Bohlinger, Jonathan Kinzel, Daniel Palenicek, Lukasz Antczak, and Jan Peters · 2025
Closest in time.
Simba: Simplicity bias for scaling up parameters in deep reinforcement learning
Hojoon Lee, Dongyoon Hwang, Donghu Kim, Hyunseung Kim, Jun Jet Tai, Kaushik Subramanian, Peter R Wurman, Jaegul Choo, Peter Stone, and Takuma Seno · 2025
Closest in time.
Mad-td: Model-augmented data stabilizes high update ratio rl, 2025
Claas A Voelcker, Marcel Hussing, Eric Eaton, Amir massoud Farahmand, and Igor Gilitschenski · 2025
Closest in time.