Fetching the paper…
Reading the bibliography…
Imitation learning has shown great promise in robotic manipulation, but the policy's execution is often unsatisfactorily slow due to commonly tardy demonstrations collected by human operators.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
An improved data anomaly detection method based on isolation forest
D. Xu, Y. Wang, Y. Meng, and Z. Zhang · 2017
Earlier work this paper cites.
hdbscan: Hierarchical density based clustering
L. McInnes, J. Healy, S. Astels, et al · 2017
Earlier work this paper cites.
Fast & accurate gaussian kernel density estimation
J. Heer · 2021
Earlier work this paper cites.
Action chunking as policy compression
L. Lai, A. Z. Huang, and S. J. Gershman · 2022
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Earlier work this paper cites.
Waypoint-based imitation learning for robotic manipulation
L. X. Shi, A. Sharma, T. Z. Zhao, and C. Finn · 2023
Earlier work this paper cites.
Multi-view masked world models for visual robotic manipulation
Y. Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel · 2023
Earlier work this paper cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Earlier work this paper cites.
π \pi 0: A vision-language-action flow model for general robot control, 2024
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Earlier work this paper cites.
Aloha unleashed: A simple recipe for robot dexterity
T. Z. Zhao, J. Tompson, D. Driess, P. Florence, K. Ghasemipour, C. Finn, and A. Wahid · 2024
Earlier work this paper cites.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Earlier work this paper cites.
Subconscious robotic imitation learning
J. Xie, Z. Wang, J. Tan, H. Lin, and X. Ma · 2024
Cited alongside, same era.
Open-television: Teleoperation with immersive active visual feedback
X. Cheng, J. Li, S. Yang, G. Yang, and X. Wang · 2024
Cited alongside, same era.
Bunny-visionpro: Real-time bimanual dexterous teleoperation for imitation learning
R. Ding, Y. Qin, J. Zhu, C. Jia, S. Yang, R. Yang, X. Qi, and X. Wang · 2024
Cited alongside, same era.
Generalizable humanoid manipulation with improved 3d diffusion policies
Y. Ze, Z. Chen, W. Wang, T. Chen, X. He, Y. Yuan, X. B. Peng, and J. Wu · 2024
Cited alongside, same era.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Consistency policy: Accelerated visuomotor policies via consistency distillation
A. Prasad, K. Lin, J. Wu, L. Zhou, and J. Bohg · 2024
Later among the works it cites.
Manicm: Real-time 3d diffusion policy via consistency model for robotic manipulation
G. Lu, Z. Gao, T. Chen, W. Dai, Z. Wang, W. Ding, and Y. Tang · 2024
Later among the works it cites.
Demogen: Synthetic demonstration generation for data-efficient visuomotor policy learning
Z. Xue, S. Deng, Z. Chen, Y. Wang, Z. Yuan, and H. Xu · 2025
Closest in time.
Accessed: 2025-2-26
https://www.figure.ai/news/helix-logistics · 2025
Closest in time.
Y. Jiang, R. Zhang, J. Wong, C. Wang, Y. Ze, H. Yin, C. Gokmen, S. Song, J. Wu, and L. Fei-Fei · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Fu, T. Z. Zhao, and C. Finn · 2024
Cited alongside, same era.
Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators
P. Wu, Y. Shentu, Z. Yi, X. Lin, and P. Abbeel · 2024
Cited alongside, same era.
Ace: A cross-platform visual-exoskeletons system for low-cost dexterous teleoperation
S. Yang, M. Liu, Y. Qin, R. Ding, J. Li, X. Cheng, R. Yang, S. Yi, and X. Wang · 2024
Cited alongside, same era.
Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation
K. Wu, C. Hou, J. Liu, Z. Che, X. Ju, Z. Yang, M. Li, Y. Zhao, Z. Xu, G. Yang, et al · 2024
Cited alongside, same era.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Cited alongside, same era.
Q. Li, Y. Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y. Deng, S. Xu, Y. Zhang, et al · 2024
Cited alongside, same era.
Closest in time.
\ \backslash pi_ { \{ 0.5 } \} : a vision-language-action model with open-world generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model
D. Qu, H. Song, Q. Chen, Y. Yao, X. Ye, Y. Ding, Z. Wang, J. Gu, B. Zhao, D. Wang, et al · 2025
Closest in time.
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model
J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al · 2025
Closest in time.
Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Huang, S. Jiang, et al · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Curating demonstrations using online experience
A. S. Chen, A. M. Lessing, Y. Liu, and C. Finn · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.