Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has driven significant progress in robotics, but its complexity and long training times remain major bottlenecks.
Deterministic policy gradient algorithms
Silver, David, Lever, Guy, Heess, Nicolas, Degris, Thomas, Wierstra, Daan, and Riedmiller, Martin · 2014
Earlier work this paper cites.
Ba, Jimmy Lei, Kiros, Jamie Ryan, and Hinton, Geoffrey E · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, Marc G, Dabney, Will, and Munos, Rémi · 2017
Earlier work this paper cites.
Emergence of locomotion behaviours in rich environments
Heess, Nicolas, Tb, Dhruva, Sriram, Srinivasan, Lemmon, Jay, Merel, Josh, Wayne, Greg, Tassa, Yuval, Erez, Tom, Wang, Ziyu, Eslami, SM, et al · 2017
Earlier work this paper cites.
Asymmetric actor critic for image-based robot learning
Pinto, Lerrel, Andrychowicz, Marcin, Welinder, Peter, Zaremba, Wojciech, and Abbeel, Pieter · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, Scott, Hoof, Herke, and Meger, David · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, Tuomas, Zhou, Aurick, Hartikainen, Kristian, Tucker, George, Ha, Sehoon, Tan, Jie, Kumar, Vikash, Zhu, Henry, Gupta, Abhishek, Abbeel, Pieter, et al · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
Hester, Todd, Vecerik, Matej, Pietquin, Olivier, Lanctot, Marc, Schaul, Tom, Piot, Bilal, Horgan, Dan, Quan, John, Sendonaris, Andrew, Osband, Ian, et al · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, Richard S and Barto, Andrew G · 2018
Earlier work this paper cites.
Learning agile and dynamic motor skills for legged robots
Hwangbo, Jemin, Lee, Joonho, Dosovitskiy, Alexey, Bellicoso, Dario, Tsounis, Vassilios, Koltun, Vladlen, and Hutter, Marco · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, Adam, Gross, Sam, Massa, Francisco, Lerer, Adam, Bradbury, James, Chanan, Gregory, Killeen, Trevor, Lin, Zeming, Gimelshein, Natalia, Antiga, Luca, et al · 2019
Earlier work this paper cites.
Isaac gym: High performance gpu-based physics simulation for robot learning, 2021
Makoviychuk, Viktor, Wawrzyniak, Lukasz, Guo, Yunrong, Lu, Michelle, Storey, Kier, Macklin, Miles, Hoeller, David, Rudin, Nikita, Allshire, Arthur, Handa, Ankur, and State, Gavriel · 2021
Cited alongside, same era.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, Antonin, Hill, Ashley, Gleave, Adam, Kanervisto, Anssi, Ernestus, Maximilian, and Dormann, Noah · 2021
Cited alongside, same era.
Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms
Huang, Shengyi, Dossa, Rousslan Fernand Julien, Ye, Chang, Braga, Jeff, Chakraborty, Dipam, Mehta, Kinal, and Araújo, João G.M · 2022
Cited alongside, same era.
Learning to walk in minutes using massively parallel deep reinforcement learning
Rudin, Nikita, Hoeller, David, Reist, Philipp, and Hutter, Marco · 2022
Cited alongside, same era.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, Pierluca, Schwarzer, Max, Nikishin, Evgenii, Bacon, Pierre-Luc, Bellemare, Marc G, and Courville, Aaron · 2023
Cited alongside, same era.
Td-mpc2: Scalable, robust world models for continuous control
Hansen, Nicklas, Su, Hao, and Wang, Xiaolong · 2024
Later among the works it cites.
Simba: Simplicity bias for scaling up parameters in deep reinforcement learning
Lee, Hojoon, Hwang, Dongyoon, Kim, Donghu, Kim, Hyunseung, Tai, Jun Jet, Subramanian, Kaushik, Wurman, Peter R, Choo, Jaegul, Stone, Peter, and Seno, Takuma · 2024
Later among the works it cites.
Normalization and effective learning rates in reinforcement learning
Lyle, Clare, Zheng, Zeyu, Khetarpal, Khimya, Martens, James, van Hasselt, Hado P, Pascanu, Razvan, and Dabney, Will · 2024
Later among the works it cites.
Reinforcement learning with action sequence for data-efficient robot learning
Seo, Younggyo and Abbeel, Pieter · 2024
Later among the works it cites.
Humanoidbench: Simulated humanoid benchmark for whole-body locomotion and manipulation
Sferrazza, Carmelo, Huang, Dun-Ming, Lin, Xingyu, Lee, Youngwoon, and Abbeel, Pieter · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mastering diverse domains through world models
Hafner, Danijar, Pasukonis, Jurgis, Ba, Jimmy, and Lillicrap, Timothy · 2023
Cited alongside, same era.
Champion-level drone racing using deep reinforcement learning
Kaufmann, Elia, Bauersfeld, Leonard, Loquercio, Antonio, Müller, Matthias, Koltun, Vladlen, and Scaramuzza, Davide · 2023
Cited alongside, same era.
Eureka: Human-level reward design via coding large language models
Ma, Yecheng Jason, Liang, William, Wang, Guanzhi, Huang, De-An, Bastani, Osbert, Jayaraman, Dinesh, Zhu, Yuke, Fan, Linxi, and Anandkumar, Anima · 2023
Cited alongside, same era.
Orbit: A unified simulation framework for interactive robot learning environments
Mittal, Mayank, Yu, Calvin, Yu, Qinxi, Liu, Jingzhou, Rudin, Nikita, Hoeller, David, Yuan, Jia Lin, Singh, Ritvik, Guo, Yunrong, Mazhar, Hammad, Mandlekar, Ajay, Babich, Buck, State, Gavriel, Hutter, Marco, and Garg, Animesh · 2023
Cited alongside, same era.
Bigger, better, faster: Human-level atari with human-level efficiency
Schwarzer, Max, Ceron, Johan Samir Obando, Courville, Aaron, Bellemare, Marc G, Agarwal, Rishabh, and Castro, Pablo Samuel · 2023
Cited alongside, same era.
Bigym: A demo-driven mobile bi-manual manipulation benchmark
Chernyadev, Nikita, Backshall, Nicholas, Ma, Xiao, Lu, Yunfan, Seo, Younggyo, and James, Stephen · 2024
Cited alongside, same era.
Simplifying deep temporal difference learning
Gallici, Matteo, Fellows, Mattie, Ellis, Benjamin, Pou, Bartomeu, Masmitja, Ivan, Foerster, Jakob Nicolaus, and Martin, Mario · 2024
Cited alongside, same era.
Towards general-purpose model-free reinforcement learning
Fujimoto, Scott, D’Oro, Pierluca, Zhang, Amy, Tian, Yuandong, and Rabbat, Michael · 2025
Closest in time.
Hyperspherical normalization for scalable deep reinforcement learning
Lee, Hojoon, Lee, Youngdo, Seno, Takuma, Kim, Donghu, Stone, Peter, and Choo, Jaegul · 2025
Closest in time.
Getting sac to work on a massive parallel simulator: An rl journey with off-policy algorithms
Raffin, Antonin · 2025
Closest in time.
Speeding up sac with massively parallel simulation
Shukla, Arth · 2025
Closest in time.
Mad-td: Model-augmented data stabilizes high update ratio rl
Voelcker, Claas A, Hussing, Marcel, Eaton, Eric, Farahmand, Amir-massoud, and Gilitschenski, Igor · 2025
Closest in time.
Zakka, Kevin, Tabanpour, Baruch, Liao, Qiayuan, Haiderbhai, Mustafa, Holt, Samuel, Luo, Jing Yuan, Allshire, Arthur, Frey, Erik, Sreenath, Koushil, Kahrs, Lueder A, et al · 2025
Closest in time.
Tdmpbc: Self-imitative reinforcement learning for humanoid robot control
Zhuang, Zifeng, Shi, Diyuan, Suo, Runze, He, Xiao, Zhang, Hongyin, Wang, Ting, Lyu, Shangke, and Wang, Donglin · 2025
Closest in time.