Td-mpc2: Scalable, robust world models for continuous control
Nicklas Hansen, Hao Su, and Xiaolong Wang · 2024
Later among the works it cites.
Dissecting deep RL with high update ratios: Combatting value divergence
Marcel Hussing, Claas A Voelcker, Igor Gilitschenski, Amir-massoud Farahmand, and Eric Eaton · 2024
Later among the works it cites.
Normalization and effective learning rates in reinforcement learning
Clare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens, Hado van Hasselt, Razvan Pascanu, and Will Dabney · 2024
Later among the works it cites.
Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control
Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Miłoś, and Marek Cygan · 2024
Later among the works it cites.
Humanoidbench: Simulated humanoid benchmark for whole-body locomotion and manipulation
Carmelo Sferrazza, Dun-Ming Huang, Xingyu Lin, Youngwoon Lee, and Pieter Abbeel · 2024
Later among the works it cites.
Gait in eight: Efficient on-robot learning for omnidirectional quadruped locomotion
Nico Bohlinger, Jonathan Kinzel, Daniel Palenicek, Lukasz Antczak, and Jan Peters · 2025
Closest in time.
Stable gradients for stable learning at scale in deep reinforcement learning
Roger Creus Castanyer, Johan Obando-Ceron, Lu Li, Pierre-Luc Bacon, Glen Berseth, Aaron Courville, and Pablo Samuel Castro · 2025
Closest in time.
Compute-optimal scaling for value-based deep RL
Original
Preston Fu, Oleh Rybkin, Zhiyuan Zhou, Michal Nauman, Pieter Abbeel, Sergey Levine, and Aviral Kumar · 2025
Closest in time.
Towards general-purpose model-free reinforcement learning
Scott Fujimoto, Pierluca D’Oro, Amy Zhang, Yuandong Tian, and Michael Rabbat · 2025
Closest in time.
A forget-and-grow strategy for deep reinforcement learning scaling in continuous control
Zilin Kang, Chenyuan Hu, Yu Luo, Zhecheng Yuan, Ruijie Zheng, and Huazhe Xu · 2025
Closest in time.
Neuroplastic expansion in deep reinforcement learning
Jiashun Liu, Johan Samir Obando Ceron, Aaron Courville, and Ling Pan · 2025
Closest in time.
ngpt: Normalized transformer with representation learning on the hypersphere
Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, and Boris Ginsburg · 2025
Closest in time.
Network sparsity unlocks the scaling potential of deep reinforcement learning
Guozheng Ma, Lu Li, Zilin Wang, Li Shen, Pierre-Luc Bacon, and Dacheng Tao · 2025
Closest in time.
Bigger, regularized, categorical: High-capacity value functions are efficient multi-task learners
Original
Michal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar, and Pieter Abbeel · 2025
Closest in time.
Scaling off-policy reinforcement learning with batch and weight normalization
Daniel Palenicek, Florian Vogt, Joe Watson, and Jan Peters · 2025
Closest in time.
Value-based deep RL scales predictably
Oleh Rybkin, Michal Nauman, Preston Fu, Charlie Victor Snell, Pieter Abbeel, Sergey Levine, and Aviral Kumar · 2025
Closest in time.
Use the online network if you can: Towards fast and stable reinforcement learning
Ahmed Hendawy, Henrik Metternich, Théo Vincent, Mahdi Kallel, Jan Peters, and Carlo D’Eramo · 2026
Closest in time.
Bridging the performance gap between target-free and target-based reinforcement learning with iterated Q-learning
Théo Vincent, Yogesh Tripathi, Tim Faust, Yaniv Oren, Jan Peters, and Carlo D’Eramo · 2026
Closest in time.