Fetching the paper…
Reading the bibliography…
Learning generalizable robot manipulation policies, especially for complex multi-fingered humanoids, remains a significant challenge.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Reinforcement learning and the reward engineering principle
D. Dewey · 2014
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
B. C. Stadie, S. Levine, and P. Abbeel · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2017
Earlier work this paper cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
R. Islam, P. Henderson, M. Gomrokchi, and D. Precup · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. Oord, and R. Munos · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
# exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. Xi Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Earlier work this paper cites.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2018
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Earlier work this paper cites.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, et al · 2018
Earlier work this paper cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen · 2018
Earlier work this paper cites.
Group normalization
Y. Wu and K. He · 2018
Earlier work this paper cites.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Cited alongside, same era.
Solving rubik’s cube with a robot hand
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Cited alongside, same era.
Benchmarking bonus-based exploration methods on the arcade learning environment
A. A. Taïga, W. Fedus, M. C. Machado, A. Courville, and M. G. Bellemare · 2019
Cited alongside, same era.
Learning quadrupedal locomotion over challenging terrain
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2020
Dextrah-rgb: Visuomotor policies to grasp anything with dexterous hands
R. Singh, A. Allshire, A. Handa, N. Ratliff, and K. Van Wyk · 2024
Later among the works it cites.
Twisting lids off with two hands
T. Lin, Z.-H. Yin, H. Qi, P. Abbeel, and J. Malik · 2024
Later among the works it cites.
Dextrah-g: Pixels-to-action dexterous arm-hand grasping with geometric fabrics
T. G. W. Lum, M. Matak, V. Makoviychuk, A. Handa, A. Allshire, T. Hermans, N. D. Ratliff, and K. Van Wyk · 2024
Later among the works it cites.
URL https://openai.com/index/learning-to-reason-with-llms
O. AI, Sep 2024 · 2024
Later among the works it cites.
Mimex: intrinsic rewards from masked input modeling
T. Lin and A. Jabri · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reward Function Design in Reinforcement Learning , pages 25–33
J. Eschmann · 2021
Cited alongside, same era.
Isaac gym: High performance gpu-based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al · 2021
Cited alongside, same era.
Advanced skills by learning locomotion and local navigation end-to-end
N. Rudin, D. Hoeller, M. Bjelonic, and M. Hutter · 2022
Cited alongside, same era.
A system for general in-hand object re-orientation
T. Chen, J. Xu, and P. Agrawal · 2022
Cited alongside, same era.
In-Hand Object Rotation via Rapid Motor Adaptation
H. Qi, A. Kumar, R. Calandra, Y. Ma, and J. Malik · 2022
Cited alongside, same era.
Champion-level drone racing using deep reinforcement learning
E. Kaufmann, L. Bauersfeld, A. Loquercio, M. Müller, V. Koltun, and D. Scaramuzza · 2023
Cited alongside, same era.
Dextreme: Transfer of agile in-hand manipulation from simulation to reality
A. Handa, A. Allshire, V. Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. Van Wyk, A. Zhurkevich, B. Sundaralingam, et al · 2023
Cited alongside, same era.
Object-centric dexterous manipulation from human motion data
Y. Chen, C. Wang, Y. Yang, and C. K. Liu · 2024
Later among the works it cites.
Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation
M. Torne, A. Simeonov, Z. Li, A. Chan, T. Chen, A. Gupta, and P. Agrawal · 2024
Later among the works it cites.
Asid: Active exploration for system identification in robotic manipulation
M. Memmel, A. Wagenmaker, C. Zhu, P. Yin, D. Fox, and A. Gupta · 2024
Later among the works it cites.
Wococo: Learning whole-body humanoid control with sequential contacts
C. Zhang, W. Xiao, T. He, and G. Shi · 2024
Later among the works it cites.
Planning-guided diffusion policy learning for generalizable contact-rich bimanual manipulation
X. Li, T. Zhao, X. Zhu, J. Wang, T. Pang, and K. Fang · 2024
Later among the works it cites.
Aloha unleashed: A simple recipe for robot dexterity
T. Z. Zhao, J. Tompson, D. Driess, P. Florence, K. Ghasemipour, C. Finn, and A. Wahid · 2024
Later among the works it cites.
Data scaling laws in imitation learning for robotic manipulation
F. Lin, Y. Hu, P. Sheng, C. Wen, J. You, and Y. Gao · 2024
Later among the works it cites.
Lessons from learning to spin” pens”
J. Wang, Y. Yuan, H. Che, H. Qi, Y. Ma, J. Malik, and X. Wang · 2024
Later among the works it cites.
Real-world humanoid locomotion with reinforcement learning
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath · 2024
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, K. Ashutosh, V. Baiyya, S. Bansal, B. Boote, et al · 2024
Later among the works it cites.
Demostart: Demonstration-led auto-curriculum applied to sim-to-real with multi-fingered robots
M. Bauza, J. E. Chen, V. Dalibard, N. Gileadi, R. Hafner, M. F. Martins, J. Moore, R. Pevceviciute, A. Laurens, D. Rao, et al · 2024
Later among the works it cites.
Visual whole-body control for legged loco-manipulation
M. Liu, Z. Chen, X. Cheng, Y. Ji, R.-Z. Qiu, R. Yang, and X. Wang · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al · 2024
Later among the works it cites.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
D.-A. et al · 2025
Closest in time.
Dexteritygen: Foundation controller for unprecedented dexterity
Z.-H. Yin, C. Wang, L. Pineda, F. Hogan, K. Bodduluri, A. Sharma, P. Lancaster, I. Prasad, M. Kalakrishnan, J. Malik, et al · 2025
Closest in time.