Fetching the paper…
Reading the bibliography…
Vision language models have played a key role in extracting meaningful features for various robotic applications.
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al. , “The” something something” video database for learning and evaluating visual common sense,” in ICCV , 2017
2017
Earlier work this paper cites.
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “Carla: An open urban driving simulator,” in CoRL , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Akbik, T. Bergmann, D. Blythe, K. Rasul, S. Schweter, and R. Vollgraf, “Flair: An easy-to-use framework for state-of-the-art nlp,” in NAACL , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Kumra, S. Joshi, and F. Sahin, “Antipodal robotic grasping using generative residual convolutional neural network,” in IROS , 2020
2020
Earlier work this paper cites.
A. Cheraghian, S. Rahman, D. Campbell, and L. Petersson, “Transductive zero-shot learning for 3d point cloud classification,” in WACV , 2020
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Earlier work this paper cites.
S. Ainetter and F. Fraundorfer, “End-to-end trainable deep neural network for robotic grasp detection and semantic segmentation from rgb,” in ICRA , 2021
2021
Earlier work this paper cites.
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in CoRL , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Zhang, Z. Guo, W. Zhang, K. Li, X. Miao, B. Cui, Y. Qiao, P. Gao, and H. Li, “Pointclip: Point cloud understanding by clip,” in CVPR , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Guzhov, F. Raue, J. Hees, and A. Dengel, “Audioclip: Extending clip to image, text and audio,” in ICASSP , 2022
2022
Earlier work this paper cites.
A. Zhan, R. Zhao, L. Pinto, P. Abbeel, and M. Laskin, “Learning visual robotic control efficiently with contrastive pre-training and data augmentation,” in IROS , 2022
2022
Earlier work this paper cites.
S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn, et al. , “Learning language-conditioned robot behavior from offline data and crowd-sourced annotation,” in CoRL , 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
D. Shah, B. Osiński, et al. , “Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,” in CoRL , 2023
2023
Cited alongside, same era.
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell, “Real-world robot learning with masked visual pre-training,” in CoRL , 2023
2023
Cited alongside, same era.
Y. J. Ma, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman, “Liv: Language-image representations and rewards for robotic control,” in ICML , 2023
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello, “Open-vocabulary panoptic segmentation with text-to-image diffusion models,” in CVPR , 2023
2023
Later among the works it cites.
J. Wang, H. Zhu, H. Guo, A. Al Mamun, C. Xiang, and T. H. Lee, “Few-shot point cloud semantic segmentation via contrastive self-supervision and multi-resolution attention,” in ICRA , 2023
2023
Later among the works it cites.
J. Liu, Y. Zhang, J.-N. Chen, J. Xiao, Y. Lu, B. A Landman, Y. Yuan, A. Yuille, Y. Tang, and Z. Zhou, “Clip-driven universal model for organ segmentation and tumor detection,” in ICCV , 2023
2023
Later among the works it cites.
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al. , “Bridgedata v2: A dataset for robot learning at scale,” in CoRL , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
R. Chen, Y. Liu, L. Kong, X. Zhu, Y. Ma, Y. Li, Y. Hou, Y. Qiao, and W. Wang, “Clip2scene: Towards label-efficient 3d scene understanding by clip,” in CVPR , 2023
2023
Cited alongside, same era.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Tao, B.-K. Bao, H. Tang, and C. Xu, “Galip: Generative adversarial clips for text-to-image synthesis,” in CVPR , 2023
2023
Cited alongside, same era.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al. , “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” ECCV , 2023
2023
Cited alongside, same era.
C. Huang, O. Mees, A. Zeng, and W. Burgard, “Visual language maps for robot navigation,” in ICRA , 2023
2023
Cited alongside, same era.
D. Auty and K. Mikolajczyk, “Learning to prompt clip for monocular depth estimation: Exploring the limits of human language,” in ICCV , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. D. Vuong, M. N. Vu, B. Huang, N. Nguyen, H. Le, T. Vo, and A. Nguyen, “Language-driven grasp detection,” in CVPR , 2024
2024
Closest in time.
2024
Closest in time.
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Sadigh, L. Guibas, and F. Xia, “Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,” in CVPR , 2024
2024
Closest in time.
P. Sermanet, T. Ding, J. Zhao, F. Xia, D. Dwibedi, K. Gopalakrishnan, C. Chan, G. Dulac-Arnold, S. Maddineni, N. J. Joshi, et al. , “Robovqa: Multimodal long-horizon reasoning for robotics,” in ICRA , 2024
2024
Closest in time.
Z. Sun, Y. Fang, T. Wu, P. Zhang, Y. Zang, S. Kong, Y. Xiong, D. Lin, and J. Wang, “Alpha-clip: A clip model focusing on wherever you want,” in CVPR , 2024
2024
Closest in time.
A. D. Vuong, M. N. Vu, H. Le, B. Huang, B. Huynh, T. Vo, A. Kugi, and A. Nguyen, “Grasp-anything: Large-scale grasp dataset from foundation models,” ICRA , 2024
2024
Closest in time.
T. Van Vo, M. N. Vu, B. Huang, T. Nguyen, N. Le, T. Vo, and A. Nguyen, “Open-vocabulary affordance detection using knowledge distillation and text-point correlation,” ICRA , 2024
2024
Closest in time.
J. Xing, L. Bauersfeld, Y. Song, C. Xing, and D. Scaramuzza, “Contrastive learning for enhancing robust scene transfer in vision-based agile flight,” in ICRA , 2024
2024
Closest in time.
X. Wang, H. Yuan, S. Zhang, D. Chen, J. Wang, Y. Zhang, Y. Shen, D. Zhao, and J. Zhou, “Videocomposer: Compositional video synthesis with motion controllability,” NeurIPS , vol. 36, 2024
2024
Closest in time.
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao, “Clip-adapter: Better vision-language models with feature adapters,” IJCV , 2024
2024
Closest in time.
N. Nguyen, M. N. Vu, B. Huang, A. Vuong, N. Le, T. Vo, and A. Nguyen, “Lightweight language-driven grasp detection using conditional consistency model,” IROS , 2024
2024
Closest in time.