Fetching the paper…
Reading the bibliography…
We present GrasMolmo, a generalizable open-vocabulary task-oriented grasping (TOG) model.
ShapeNet: An Information-Rich 3D Model Repository
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, J. Xiao, L. Yi, and F. Yu · 2015
Earlier work this paper cites.
Semantically-Enriched 3D Models for Common-sense Knowledge
M. Savva, A. X. Chang, and P. Hanrahan · 2015
Earlier work this paper cites.
Task-oriented grasping with semantic and geometric scene understanding
R. Detry, J. Papon, and L. Matthies · 2017
Earlier work this paper cites.
Learning task-oriented grasping for tool manipulation from simulated self-supervision
K. Fang, Y. Zhu, A. Garg, A. Kurenkov, V. Mehta, L. Fei-Fei, and S. Savarese · 2018
Earlier work this paper cites.
Affordancenet: An end-to-end deep learning approach for object affordance detection
T.-T. Do, A. Nguyen, and I. Reid · 2018
Earlier work this paper cites.
Dense object nets: Learning dense visual object descriptors by and for robotic manipulation
P. Florence, L. Manuelli, and R. Tedrake · 2018
Earlier work this paper cites.
kpam: Keypoint affordances for category-level robotic manipulation
L. Manuelli, W. Gao, P. Florence, and R. Tedrake · 2019
Earlier work this paper cites.
Same object, different grasps: Data and semantic knowledge for task-oriented grasping
A. Murali, W. Liu, K. Marino, S. Chernova, and A. Gupta · 2020
Earlier work this paper cites.
6-dof grasping for target-driven object manipulation in clutter
A. Murali, A. Mousavian, C. Eppner, C. Paxton, and D. Fox · 2020
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2020
Earlier work this paper cites.
Keto: Learning keypoint representations for tool manipulation
Z. Qin, K. Fang, Y. Zhu, L. Fei-Fei, and S. Savarese · 2020
Earlier work this paper cites.
Deepgmr: Learning latent gaussian mixture models for registration
W. Yuan, B. Eckart, K. Kim, V. Jampani, D. Fox, and J. Kautz · 2020
Earlier work this paper cites.
Contact-graspnet: Efficient 6-dof grasp generation in cluttered scenes
M. Sundermeyer, A. Mousavian, R. Triebel, and D. Fox · 2021
Earlier work this paper cites.
Synergies between affordance and geometry: 6-dof grasp detection via implicit representations
Z. Jiang, Y. Zhu, M. Svetlik, K. Fang, and Y. Zhu · 2021
Earlier work this paper cites.
Where2act: From pixels to actions for articulated 3d objects
K. Mo, L. J. Guibas, M. Mukadam, A. Gupta, and S. Tulsiani · 2021
Earlier work this paper cites.
Acronym: A large-scale grasp dataset based on simulation
C. Eppner, A. Mousavian, and D. Fox · 2021
Earlier work this paper cites.
Structformer: Learning spatial structure for language-guided semantic rearrangement of novel objects
W. Liu, C. Paxton, T. Hermans, and D. Fox · 2022
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Cited alongside, same era.
Task-oriented grasp prediction with visual-language inputs
C. Tang, D. Huang, L. Meng, W. Liu, and H. Zhang · 2023
Cited alongside, same era.
Graspgpt: Leveraging semantic knowledge from a large language model for task-oriented grasping
C. Tang, D. Huang, W. Ge, W. Liu, and H. Zhang · 2023
Cited alongside, same era.
Language embedded radiance fields for zero-shot task-oriented grasping
J. Kerr, C. M. Kim, K. Goldberg, L. Y. Chen, S. Sharma, A. Rashid, and A. Kanazawa · 2023
Cited alongside, same era.
Language-guided robot grasping: Clip-based referring grasp synthesis in clutter
G. Tziafas, X. Yucheng, A. Goel, M. Kasaei, Z. Li, and H. Kasaei · 2023
Cited alongside, same era.
Task-oriented grasp prediction with visual-language inputs
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Later among the works it cites.
π \pi 0 {}_{\mbox{0}} : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky · 2024
Later among the works it cites.
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
F. Liu, K. Fang, P. Abbeel, and S. Levine · 2024
Later among the works it cites.
Manipulate-anything: Automating real-world robots using vision-language models
J. Duan, W. Yuan, W. Pumacay, Y. R. Wang, K. Ehsani, D. Fox, and R. Krishna · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Tang, D. Huang, L. Meng, W. Liu, and H. Zhang · 2023
Cited alongside, same era.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
RT-2: vision-language-action models transfer web knowledge to robotic control
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. T. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T. E. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. J. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, K. A. Dubey, D. Driess, T. Ding, K. M. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Cited alongside, same era.
M2t2: Multi-task masked transformer for object-centric pick and place
W. Yuan, A. Murali, A. Mousavian, and D. Fox · 2023
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Cited alongside, same era.
X. Li, C. Mata, J. Park, K. Kahatapitiya, Y. S. Jang, J. Shang, K. Ranasinghe, R. Burgert, M. Cai, Y. J. Lee, et al · 2024
Later among the works it cites.
Feedback-guided autonomous driving
J. Zhang, Z. Huang, A. Ray, and E. Ohn-Bar · 2024
Later among the works it cites.
K. Ehsani, T. Gupta, R. Hendrix, J. Salvador, L. Weihs, K.-H. Zeng, K. P. Singh, Y. Kim, W. Han, A. Herrasti, R. Krishna, D. Schwenk, E. VanderBilt, and A. Kembhavi · 2024
Later among the works it cites.
Robopoint: A vision-language model for spatial affordance prediction for robotics
W. Yuan, J. Duan, V. Blukis, W. Pumacay, R. Krishna, A. Murali, A. Mousavian, and D. Fox · 2024
Later among the works it cites.
Sat: Spatial aptitude training for multimodal language models
A. Ray, J. Duan, R. Tan, D. Bashkirova, R. Hendrix, K. Ehsani, A. Kembhavi, B. A. Plummer, R. Krishna, K.-H. Zeng, et al · 2024
Later among the works it cites.
Poliformer: Scaling on-policy rl with transformers results in masterful navigators, 2024
K.-H. Zeng, Z. Zhang, K. Ehsani, R. Hendrix, J. Salvador, A. Herrasti, R. Girshick, A. Kembhavi, and L. Weihs · 2024
Later among the works it cites.
What do we learn from a large-scale study of pre-trained visual representations in sim and real environments?
S. Silwal, K. Yadav, T. Wu, J. Vakil, A. Majumdar, S. Arnaud, C. Chen, V.-P. Berges, D. Batra, A. Rajeswaran, et al · 2024
Later among the works it cites.
scene_synthesizer: A python library for procedural scene generation in robot manipulation
C. Eppner, A. Murali, C. Garrett, R. O’Flaherty, T. Hermans, W. Yang, and D. Fox · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer · 2024
Later among the works it cites.
GR00T N1: an open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, Linxi, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. LLontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu · 2025
Closest in time.
Physbench: Benchmarking and enhancing vision-language models for physical world understanding
W. Chow, J. Mao, B. Li, D. Seita, V. Guizilini, and Y. Wang · 2025
Closest in time.
Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics, 2025
C. H. Song, V. Blukis, J. Tremblay, S. Tyree, Y. Su, and S. Birchfield · 2025
Closest in time.