Fetching the paper…
Reading the bibliography…
Despite recent progress in general purpose robotics, robot policies still lag far behind basic human capabilities in the real world.
3d hand shape and pose from images in the wild
A. Boukhayma, R. A. de Bem, and P. H. S. Torr · 1902
Earlier work this paper cites.
Pushing the envelope for rgb-based dense 3d hand pose estimation via neural rendering
S. Baek, K. I. Kim, and T. Kim · 1904
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning, 2019
OpenAI · 1912
Earlier work this paper cites.
Toward automatic robot instruction from perception—mapping human grasps to manipulator grasps
S.-R. Kang and K. Ikeuchi · 1994
Earlier work this paper cites.
Language models are few-shot learners
OpenAI · 2005
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Imitation learning: A survey of learning methods
A. Hussein, M. M. Gaber, E. Elyan, and C. Jayne · 2017
Earlier work this paper cites.
The ”something something” video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic · 2017
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Earlier work this paper cites.
Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity
A. Mandlekar, J. Booher, M. Spero, A. Tung, A. Gupta, Y. Zhu, A. Garg, S. Savarese, and L. Fei-Fei · 2019
Earlier work this paper cites.
End-to-end hand mesh recovery from a monocular rgb image
X. Zhang, Q. Li, H. Mo, W. Zhang, and W. Zheng · 2019
Earlier work this paper cites.
Understanding human hands in contact at internet scale
D. Shan, J. Geng, M. Shu, and D. Fouhey · 2020
Earlier work this paper cites.
Zero-shot text-to-image generation, 2021
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Earlier work this paper cites.
Dexycb: A benchmark for capturing hand grasping of objects
Y.-W. Chao, W. Yang, Y. Xiang, P. Molchanov, A. Handa, J. Tremblay, Y. S. Narang, K. Van Wyk, U. Iqbal, S. Birchfield, et al · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2022
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Robust speech recognition via large-scale weak supervision, 2022
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Earlier work this paper cites.
Human-to-robot imitation in the wild, 2022
S. Bahl, A. Gupta, and D. Pathak · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
BC-Z: zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
Dexmv: Imitation learning for dexterous manipulation from human videos
Y. Qin, Y.-H. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang · 2022
Cited alongside, same era.
Embodied hands: Modeling and capturing hands and bodies together
J. Romero, D. Tzionas, and M. J. Black · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, V. Jampani, and R. Rombach · 2023
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers, 2023
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Aloha unleashed: A simple recipe for robot dexterity, 2024
T. Z. Zhao, J. Tompson, D. Driess, P. Florence, K. Ghasemipour, C. Finn, and A. Wahid · 2024
Later among the works it cites.
Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators, 2024
P. Wu, Y. Shentu, Z. Yi, X. Lin, and P. Abbeel · 2024
Later among the works it cites.
Semantic constraints to represent common sense required in household actions for multimodal learning-from-observation robot
K. Ikeuchi, K. Minamizawa, K. Harada, A. Yamaguchi, and S. Kagami · 2024
Later among the works it cites.
https://www.meta.com/quest/ , 2024
Meta quest · 2024
Later among the works it cites.
https://www.apple.com/apple-vision-pro/ , 2024
Apple vision pro · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Cited alongside, same era.
Project aria: A new tool for egocentric multi-modal ai research, 2023
J. Engel, K. Somasundaram, M. Goesele, A. Sun, A. Gamino, A. Turner, A. Talattof, A. Yuan, B. Souti, B. Meredith, C. Peng, C. Sweeney, C. Wilson, D. Barnes, D. DeTone, D. Caruso, D. Valleroy, D. Ginjupalli, D. Frost, E. Miller, E. Mueggler, E. Oleinik, F. Zhang, G. Somasundaram, G. Solaira, H. Lanaras, H. Howard-Jenkins, H. Tang, H. J. Kim, J. Rivera, J. Luo, J. Dong, J. Straub, K. Bailey, K. Eckenhoff, L. Ma, L. Pesqueira, M. Schwesinger, M. Monge, N. Yang, N. Charron, N. Raina, O. Parkhi, P. Borschowa, P. Moulon, P. Gupta, R. Mur-Artal, R. Pennington, S. Kulkarni, S. Miglani, S. Gondi, S. Solanki, S. Diener, S. Cheng, S. Green, S. Saarinen, S. Patra, T. Mourikis, T. Whelan, T. Singh, V. Balntas, V. Baiyya, W. Dreewes, X. Pan, Y. Lou, Y. Zhao, Y. Mansour, Y. Zou, Z. Lv, Z. Wang, M. Yan, C. Ren, R. D. Nardi, and R. Newcombe · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Cited alongside, same era.
Designing anthropomorphic soft hands through interaction
P. Mannam, K. Shaw, D. Bauer, J. Oh, D. Pathak, and N. Pollard · 2023
Cited alongside, same era.
Affordances from human videos as a versatile representation for robotics
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak · 2023
Cited alongside, same era.
Reconstructing hands in 3d with transformers
G. Pavlakos, D. Shan, I. Radosavovic, A. Kanazawa, D. Fouhey, and J. Malik · 2023
Cited alongside, same era.
Emergent correspondence from image diffusion, 2023
L. Tang, M. Jia, Q. Wang, C. P. Phoo, and B. Hariharan · 2023
Cited alongside, same era.
OpenAI · 2024
Cited alongside, same era.
https://store.steampowered.com/app/250820/SteamVR/ , 2024
Steamvr · 2024
Later among the works it cites.
https://www.movella.com/products/xsens , 2024
Movella xsens · 2024
Later among the works it cites.
https://www.manus-meta.com , 2024
Manusmetagloves · 2024
Later among the works it cites.
https://www.rokoko.com , 2024
Rokoko · 2024
Later among the works it cites.
R+x: Retrieval and execution from everyday human videos, 2024
G. Papagiannis, N. D. Palo, P. Vitiello, and E. Johns · 2024
Later among the works it cites.
Hand-object interaction pretraining from videos, 2024
H. G. Singh, A. Loquercio, C. Sferrazza, J. Wu, H. Qi, P. Abbeel, and J. Malik · 2024
Later among the works it cites.
Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024
N. Karaev, I. Makarov, J. Wang, N. Neverova, A. Vedaldi, and C. Rupprecht · 2024
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2024
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang · 2024
Later among the works it cites.
Baku: An efficient transformer for multi-task policy learning, 2024
S. Haldar, Z. Peng, and L. Pinto · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model, 2024
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos, 2024
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer · 2024
Later among the works it cites.
Point policy: Unifying observations and actions with key points for robot manipulation, 2025
S. Haldar and L. Pinto · 2025
Closest in time.
Zeromimic: Distilling robotic manipulation skills from web videos, 2025
J. Shi, Z. Zhao, T. Wang, I. Pedroza, A. Luo, J. Wang, J. Ma, and D. Jayaraman · 2025
Closest in time.
Phantom: Training robots without robots using only human videos, 2025
M. Lepert, J. Fang, and J. Bohg · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization, 2025
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky · 2025
Closest in time.
Depth pro: Sharp monocular metric depth in less than a second
A. Bochkovskii, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. R. Richter, and V. Koltun · 2025
Closest in time.