Fetching the paper…
Reading the bibliography…
Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data.
Scaling egocentric vision: The epic-kitchens dataset
Damen, D.; Doughty, H.; Farinella, G. M.; Fidler, S.; Furnari, A.; Kazakos, E.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; et al. 2018 · 2018
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. 2022 · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Lipman, Y.; Chen, R. T.; Ben-Hamu, H.; Nickel, M.; and Le, M. 2022 · 2022
Earlier work this paper cites.
Rectified flow: A marginal preserving approach to optimal transport
Liu, Q. 2022 · 2022
Earlier work this paper cites.
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Liu, Y.; Liu, Y.; Jiang, C.; Lyu, K.; Wan, W.; Shen, H.; Liang, B.; Fu, Z.; Wang, H.; and Yi, L. 2022 · 2022
Earlier work this paper cites.
All are worth words: A vit backbone for diffusion models
Bao, F.; Nie, S.; Xue, K.; Cao, Y.; Li, C.; Su, H.; and Zhu, J. 2023 · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; et al. 2023 · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; and Song, S. 2023 · 2023
Earlier work this paper cites.
Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot
Fang, H.-S.; Fang, H.; Tang, Z.; Liu, J.; Wang, C.; Wang, J.; Zhu, H.; and Lu, C. 2023 · 2023
Earlier work this paper cites.
Dinov2: Learning robust visual features without supervision
Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. 2023 · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
Peebles, W.; and Xie, S. 2023 · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training
Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023 · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T. Z.; Kumar, V.; Levine, S.; and Finn, C. 2023 · 2023
Earlier work this paper cites.
Aloha 2: An enhanced low-cost hardware for bimanual teleoperation
Aldaco, J.; Armstrong, T.; Baruch, R.; Bingham, J.; Chan, S.; Draper, K.; Dwibedi, D.; Finn, C.; Florence, P.; Goodrich, S.; et al. 2024 · 2024
Cited alongside, same era.
π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control
Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. 2024 · 2024
Cited alongside, same era.
Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots
Chi, C.; Xu, Z.; Pan, C.; Cousineau, E.; Burchfiel, B.; Feng, S.; Tedrake, R.; and Song, S. 2024 · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. 2024 · 2024
Cited alongside, same era.
Flow as the Cross-Domain Manipulation Interface
Xu, M.; Zhang, Z.; Chi, C.; and Song, S. 2024 · 2024
Later among the works it cites.
3d diffusion policy
Ze, Y.; Zhang, G.; Zhang, K.; Hu, C.; Wang, M.; and Xu, H. 2024 · 2024
Later among the works it cites.
3d-vla: A 3d vision-language-action generative world model
Zhen, H.; Qiu, X.; Chen, P.; Yang, J.; Yan, X.; Du, Y.; Hong, Y.; and Gan, C. 2024 · 2024
Later among the works it cites.
Hot3d: Hand and object tracking in 3d from egocentric multi-view videos
Banerjee, P.; Shkodrani, S.; Moulon, P.; Hampali, S.; Han, S.; Zhang, F.; Zhang, L.; Fountain, J.; Miller, E.; Basol, S.; et al. 2025 · 2025
Closest in time.
Bu, Q.; Cai, J.; Chen, L.; Cui, X.; Ding, Y.; Feng, S.; Gao, S.; He, X.; Hu, X.; Huang, X.; et al. 2025 · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Grauman, K.; Westbury, A.; Torresani, L.; Kitani, K.; Malik, J.; Afouras, T.; Ashutosh, K.; Baiyya, V.; Bansal, S.; Boote, B.; et al. 2024 · 2024
Cited alongside, same era.
Egomimic: Scaling imitation learning via egocentric video
Kareer, S.; Patel, D.; Punamiya, R.; Mathur, P.; Cheng, S.; Wang, C.; Hoffman, J.; and Xu, D. 2024 · 2024
Cited alongside, same era.
Droid: A large-scale in-the-wild robot manipulation dataset
Khazatsky, A.; Pertsch, K.; Nair, S.; Balakrishna, A.; Dasari, S.; Karamcheti, S.; Nasiriany, S.; Srirama, M. K.; Chen, L. Y.; Ellis, K.; et al. 2024 · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; et al. 2024 · 2024
Cited alongside, same era.
Rdt-1b: a diffusion foundation model for bimanual manipulation
Liu, S.; Wu, L.; Li, B.; Tan, H.; Chen, H.; Wang, Z.; Xu, K.; Su, H.; and Zhu, J. 2024 · 2024
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
O’Neill, A.; Rehman, A.; Maddukuri, A.; Gupta, A.; Padalkar, A.; Lee, A.; Pooley, A.; Gupta, A.; Mandlekar, A.; Jain, A.; et al. 2024 · 2024
Cited alongside, same era.
Octo: An open-source generalist robot policy
Team, O. M.; Ghosh, D.; Walke, H.; Pertsch, K.; Black, K.; Mees, O.; Dasari, S.; Hejna, J.; Kreiman, T.; Xu, C.; et al. 2024 · 2024
Cited alongside, same era.
Dexcap: Scalable and portable mocap data collection system for dexterous manipulation
Wang, C.; Shi, H.; Wang, W.; Zhang, R.; Fei-Fei, L.; and Liu, C. K. 2024 · 2024
Cited alongside, same era.
Closest in time.
Chen, T.; Chen, Z.; Chen, B.; Cai, Z.; Liu, Y.; Liang, Q.; Li, Z.; Lin, X.; Ge, Y.; Gu, Z.; et al. 2025 · 2025
Closest in time.
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
Hoque, R.; Huang, P.; Yoon, D. J.; Sivapurapu, M.; and Zhang, J. 2025 · 2025
Closest in time.
π 0.5 \pi_{0.5} : A Vision-Language-Action Model with Open-World Generalization
Intelligence, P.; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; et al. 2025 · 2025
Closest in time.
H2R: Learning Human-to-Robot Imitation with Data Augmentation
Li, Z.; Wang, J.; Chen, H.; Liu, M.; et al. 2025 · 2025
Closest in time.
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model
Liu, J.; Chen, H.; An, P.; Liu, Z.; Zhang, R.; Gu, C.; Li, X.; Guo, Z.; Chen, S.; Liu, M.; et al. 2025 · 2025
Closest in time.
Qiu, R.-Z.; Yang, S.; Cheng, X.; Chawla, C.; Li, J.; He, T.; Yan, G.; Yoon, D. J.; Hoque, R.; Paulsen, L.; et al. 2025 · 2025
Closest in time.
Human2Robot: Learning Robot Actions from Paired Human-Robot Videos
Xie, S.; Cao, H.; Weng, Z.; Xing, Z.; Shen, S.; Leng, J.; Qiu, X.; Fu, Y.; Wu, Z.; and Jiang, Y.-G. 2025 · 2025
Closest in time.
DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation
Xu, M.; Zhang, H.; Hou, Y.; Xu, Z.; Fan, L.; Veloso, M.; and Song, S. 2025 · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Zhao, Q.; Lu, Y.; Kim, M. J.; Fu, Z.; Zhang, Z.; Wu, Y.; Li, Z.; Ma, Q.; Han, S.; Finn, C.; et al. 2025 · 2025
Closest in time.