Fetching the paper…
Reading the bibliography…
Real robot data collection for imitation learning has led to significant advancements in robotic manipulation.
On characterizing and computing three-and four-finger force-closure grasps of polyhedral objects
J. Ponce, S. Sullivan, J.-D. Boissonnat, and J.-P. Merlet · 1993
Earlier work this paper cites.
On computing four-finger equilibrium and force-closure grasps of polyhedral objects
J. Ponce, S. Sullivan, A. Sudsang, J.-D. Boissonnat, and J.-P. Merlet · 1997
Earlier work this paper cites.
Distance between a point and a convex cone in n n -dimensional space: Computation and applications
Y. Zheng and C.-M. Chew · 2009
Earlier work this paper cites.
From caging to grasping
A. Rodriguez, M. T. Mason, and S. Ferry · 2012
Earlier work this paper cites.
On the synthesis of feasible and prehensile robotic grasps
C. Rosales, R. Suárez, M. Gabiccini, and A. Bicchi · 2012
Earlier work this paper cites.
On the manipulability ellipsoids of underactuated robotic hands with compliance
D. Prattichizzo, M. Malvezzi, M. Gabiccini, and A. Bicchi · 2012
Earlier work this paper cites.
Delving into egocentric actions
Y. Li, Z. Ye, and J. M. Rehg · 2015
Earlier work this paper cites.
Embodied hands: Modeling and capturing hands and bodies together
J. Romero, D. Tzionas, and M. J. Black · 2017
Earlier work this paper cites.
Synthesis and optimization of force closure grasps via sequential semidefinite programming
H. Dai, A. Majumdar, and R. Tedrake · 2018
Earlier work this paper cites.
Roboturk: A crowdsourcing platform for robotic skill learning through imitation
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, et al · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Actor and observer: Joint modeling of first and third-person videos
G. A. Sigurdsson, A. K. Gupta, C. Schmid, A. Farhadi, and A. Karteek · 2018
Earlier work this paper cites.
On the effectiveness of task granularity for transfer learning
F. Mahdisoltani, G. Berger, W. Gharbieh, D. Fleet, and R. Memisevic · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Earlier work this paper cites.
Contactgrasp: Functional multi-finger grasp synthesis from contact
S. Brahmbhatt, A. Handa, J. Hays, and D. Fox · 2019
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
Learning dexterous in-hand manipulation
O. M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2020
Earlier work this paper cites.
Deep dynamics models for learning dexterous manipulation
A. Nagabandi, K. Konolige, S. Levine, and V. Kumar · 2020
Earlier work this paper cites.
Ganhand: Predicting human grasp affordances in multi-object scenes
E. Corona, A. Pumarola, G. Alenya, F. Moreno-Noguer, and G. Rogez · 2020
Earlier work this paper cites.
Unigrasp: Learning a unified model to grasp with multifingered robotic hands
L. Shao, F. Ferreira, M. Jorda, V. Nambiar, J. Luo, E. Solowjow, J. A. Ojea, O. Khatib, and J. Bohg · 2020
Earlier work this paper cites.
On the continuity of rotation representations in neural networks, 2020
Y. Zhou, C. Barnes, J. Lu, J. Yang, and H. Li · 2020
Earlier work this paper cites.
The vicarios virtual reality interface for remote robotic teleoperation: Teleporting for intuitive tele-manipulation
A. Naceri, D. Mazzanti, J. Bimbo, Y. T. Tefera, D. Prattichizzo, D. G. Caldwell, L. S. Mattos, and N. Deshpande · 2021
Earlier work this paper cites.
Hand-object contact consistency reasoning for human grasps generation
H. Jiang, S. Liu, J. Wang, and X. Wang · 2021
Earlier work this paper cites.
Cpf: Learning a contact potential field to model the hand-object interaction
L. Yang, X. Zhan, K. Li, W. Xu, J. Li, and C. Lu · 2021
Earlier work this paper cites.
Learning diverse and physically feasible dexterous grasps with generative model and bilevel optimization
A. Wu, M. Guo, and C. K. Liu · 2022
Earlier work this paper cites.
Grasp’d: Differentiable contact-rich grasp synthesis for multi-fingered hands
D. Turpin, L. Wang, E. Heiden, Y.-C. Chen, M. Macklin, S. Tsogkas, S. Dickinson, and A. Garg · 2022
Earlier work this paper cites.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
D. Damen, H. Doughty, G. M. Farinella, A. Furnari, E. Kazakos, J. Ma, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation, 2022
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Joint hand motion and interaction hotspots prediction from egocentric videos
S. Liu, S. Tripathi, S. Majumdar, and X. Wang · 2022
Cited alongside, same era.
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Y. Liu, Y. Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi · 2022
Cited alongside, same era.
Bunny-visionpro: Real-time bimanual dexterous teleoperation for imitation learning
R. Ding, Y. Qin, J. Zhu, C. Jia, S. Yang, R. Yang, X. Qi, and X. Wang · 2024
Later among the works it cites.
Gaze-guided hand-object interaction synthesis: Benchmark and method
J. Tian, L. Yang, R. Ji, Y. Ma, L. Xu, J. Yu, Y. Shi, and J. Wang · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine · 2024
Later among the works it cites.
OpenVLA: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Later among the works it cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Open x-embodiment: Robotic learning datasets and RT-x models
Q. Vuong, S. Levine, H. R. Walke, K. Pertsch, and A. S. et al · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Cited alongside, same era.
Gpt-4 technical report
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
What does a platypus look like? generating customized prompts for zero-shot image classification
S. Pratt, I. Covert, R. Liu, and A. Farhadi · 2023
Cited alongside, same era.
Open-vocabulary object detection upon frozen vision and language models
W. Kuo, Y. Cui, X. Gu, A. Piergiovanni, and A. Angelova · 2023
Cited alongside, same era.
Kosmos-2.5: A multimodal literate model
T. Lv, Y. Huang, J. Chen, Y. Zhao, Y. Jia, L. Cui, S. Ma, Y. Chang, S. Huang, W. Wang, et al · 2023
Cited alongside, same era.
Mimicgen: A data generation system for scalable robot learning using human demonstrations
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y. Narang, L. Fan, Y. Zhu, and D. Fox · 2023
Cited alongside, same era.
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Later among the works it cites.
Isaac sim: Advanced simulation for robotics development, 2024
NVIDIA · 2024
Later among the works it cites.
Egomimic: Scaling imitation learning via egocentric video, 2024
S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu · 2024
Later among the works it cites.
Vila: On pre-training for visual language models
J. Lin, H. Yin, W. Ping, P. Molchanov, M. Shoeybi, and S. Han · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, et al · 2024
Later among the works it cites.
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
J. Wen, Y. Zhu, J. Li, M. Zhu, K. Wu, Z. Xu, R. Cheng, C. Shen, Y. Peng, F. Feng, et al · 2024
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, K. Ashutosh, V. Baiyya, S. Bansal, B. Boote, et al · 2024
Later among the works it cites.
Spatiotemporal predictive pre-training for robotic motor control, 2024
J. Yang, B. Liu, J. Fu, B. Pan, G. Wu, and L. Wang · 2024
Later among the works it cites.
Learning manipulation by predicting interaction
J. Zeng, Q. Bu, B. Wang, W. Xia, L. Chen, H. Dong, H. Song, D. Wang, D. Hu, P. Luo, et al · 2024
Later among the works it cites.
Latent action pretraining from videos, 2024
S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y.-W. Chao, B. Y. Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo · 2024
Later among the works it cites.
Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers
W. Lirui, C. Xinlei, Z. Jialiang, and H. Kaiming · 2024
Later among the works it cites.
Evaluating real-world robot manipulation policies in simulation, 2024
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao · 2024
Later among the works it cites.
H1, 2024
U. Robotics · 2024
Later among the works it cites.
The dexterous hands, 2024
I. Robots · 2024
Later among the works it cites.
Taco: Benchmarking generalizable bimanual tool-action-object understanding
Y. Liu, H. Yang, X. Si, L. Liu, Z. Li, Y. Zhang, Y. Liu, and L. Yi · 2024
Later among the works it cites.
Introducing hot3d: An egocentric dataset for 3d hand and object tracking
P. Banerjee, S. Shkodrani, P. Moulon, S. Hampali, F. Zhang, J. Fountain, E. Miller, S. Basol, R. Newcombe, R. Wang, J. J. Engel, and T. Hodan · 2024
Later among the works it cites.
Nvidia omniverse: A platform for virtual collaboration and real-time simulation, 2024
NVIDIA · 2024
Later among the works it cites.
Humanoid policy human policy, 2025
R.-Z. Qiu, S. Yang, X. Cheng, C. Chawla, J. Li, T. He, G. Yan, D. J. Yoon, R. Hoque, L. Paulsen, G. Yang, J. Zhang, S. Yi, G. Shi, and X. Wang · 2025
Closest in time.
Myvlm: Personalizing vlms for user-specific queries
Y. Alaluf, E. Richardson, S. Tulyakov, K. Aberman, and D. Cohen-Or · 2025
Closest in time.
Lita: Language instructed temporal-localization assistant
D.-A. Huang, S. Liao, S. Radhakrishnan, H. Yin, P. Molchanov, Z. Yu, and J. Kautz · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization, 2025
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky · 2025
Closest in time.
Nvila: Efficient frontier visual language models, 2025
Z. Liu, L. Zhu, B. Shi, Z. Zhang, Y. Lou, S. Yang, H. Xi, S. Cao, Y. Gu, D. Li, X. Li, Y. Fang, Y. Chen, C.-Y. Hsieh, D.-A. Huang, A.-C. Cheng, V. Nath, J. Hu, S. Liu, R. Krishna, D. Xu, X. Wang, P. Molchanov, J. Kautz, H. Yin, S. Han, and Y. Lu · 2025
Closest in time.