Fetching the paper…
Reading the bibliography…
Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding, however they are still struggling with spatial understanding which is the foundation of Embodied AI.
Spatial cognition: The role of landmark, route, and survey knowledge in human and robot navigation
Steffen Werner, Bernd Krieg-Brückner, Hanspeter A. Mallot, Karin Schweizer, and Christian Freksa · 1997
Earlier work this paper cites.
Indoor navigation of a wheeled mobile robot along visual routes
Guillaume Le Blanc, Youcef Mezouar, and Philippe Martinet · 2005
Earlier work this paper cites.
Kinectfusion: real-time 3d reconstruction and interaction using a moving depth camera
Shahram Izadi, David Kim, Otmar Hilliges, David Molyneaux, Richard A. Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Dustin Freeman, Andrew J. Davison, and Andrew William Fitzgibbon · 2011
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Computational photography with plenoptic camera and light field capture: tutorial
Edmund Y. Lam · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Dense 3d reconstruction combining depth and rgb information
Hailong Pan, Tao Guan, Yawei Luo, Liya Duan, Yuan Tian, Liu Yi, Yizhu Zhao, and Junqing Yu · 2016
Earlier work this paper cites.
Joint 2d-3d-semantic data for indoor scene understanding
Iro Armeni, Sasha Sax, Amir Zamir, and Silvio Savarese · 2017
Earlier work this paper cites.
Spatial as deep: Spatial cnn for traffic scene understanding
Xingang Pan, Jianping Shi, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2017
Earlier work this paper cites.
Sparsity invariant cnns
Jonas Uhrig, N. Schneider, Lukas Schneider, Uwe Franke, Thomas Brox, and Andreas Geiger · 2017
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, K. Batmanghelich, and Dacheng Tao · 2018
Earlier work this paper cites.
A comprehensive study for robot navigation techniques
Faiza Gul, Wan Rahiman, and Syed Sahal Nazli Alhady · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning · 2019
Earlier work this paper cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2019
Earlier work this paper cites.
Semantic information for robot navigation: A survey
Jonathan Crespo, José Carlos Castillo, Óscar Martínez Mozos, and Ramón Barber · 2020
Earlier work this paper cites.
Graspnet-1billion: A large-scale benchmark for general object grasping
Haoshu Fang, Chenxi Wang, Minghao Gou, and Cewu Lu · 2020
Earlier work this paper cites.
A review of techniques for 3d reconstruction of indoor environments
Zhizhong Kang, Juntao Yang, Zhou Yang, and Sai Cheng · 2020
Earlier work this paper cites.
Scanqa: 3d question answering for spatial scene understanding
Daich Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoaki Kawanabe · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox · 2021
Earlier work this paper cites.
3d spatial measurement for model reconstruction: A review
Wendy Flores-Fuentes, Gabriel Trujillo-Hernández, Iván Y. Alba-Corpus, Julio César Rodríguez-Quiñonez, Jesús E. Mirada-Vega, Daniel Hernández-Balbuena, Fabian Natanael Murrieta-Rico, and Oleg Yu. Sergiyenko · 2022
Cited alongside, same era.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi (Jim) Fan · 2022
Cited alongside, same era.
Learning robotic navigation from experience: principles, methods and recent results
Sergey Levine and Dhruv Shah · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Cited alongside, same era.
Ferret: Refer and ground anything anywhere at any granularity
Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, and Yinfei Yang · 2023
Later among the works it cites.
Osprey: Pixel understanding with visual instruction tuning
Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, and Jianke Zhu · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Renrui Zhang, Ziyao Zeng, and Ziyu Guo · 2022
Cited alongside, same era.
Learning to prompt clip for monocular depth estimation: Exploring the limits of human language
Dylan Auty and Krystian Mikolajczyk · 2023
Cited alongside, same era.
Zoedepth: Zero-shot transfer by combining relative and metric depth
S. Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Muller · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Zhen fei Yin, Jiong Wang, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Lu Sheng, Lei Bai, Xiaoshui Huang, Zhiyong Wang, Wanli Ouyang, and Jing Shao · 2023
Cited alongside, same era.
Finetuning offline world models in the real world
Yunhai Feng, Nicklas Hansen, Ziyan Xiong, Chandramouli Rajagopalan, and Xiaolong Wang · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, and Rongrong Ji · 2023
Cited alongside, same era.
Physically grounded vision-language models for robotic manipulation
Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao, Jiajun Wu, Brian Ichter, Anirudha Majumdar, and Dorsa Sadigh · 2023
Cited alongside, same era.
Bo Zhao, Boya Wu, Muyang He, and Tiejun Huang · 2023
Later among the works it cites.
Segment everything everywhere all at once
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, and Yong Jae Lee · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
Closest in time.
Language-based depth hints for monocular depth estimation
Dylan Auty and Krystian Mikolajczyk · 2024
Closest in time.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Danny Driess, Pete Florence, Dorsa Sadigh, Leonidas Guibas, and Fei Xia · 2024
Closest in time.
Language-image models with 3d understanding
Jang Hyun Cho, B. Ivanovic, Yulong Cao, Edward Schmerling, Yue Wang, Xinshuo Weng, Boyi Li, Yurong You, Philipp Krahenbuhl, Yan Wang, and Marco Pavone · 2024
Closest in time.
Efficient multimodal learning from data-centric perspective
Muyang He, Yexin Liu, Boya Wu, Jianhao Yuan, Yueze Wang, Tiejun Huang, and Bo Zhao · 2024
Closest in time.
Visual hallucinations of multi-modal large language models
Wen Huang, Hongbin Liu, Minxin Guo, and Neil Zhenqiang Gong · 2024
Closest in time.
Efficient multimodal large language models: A survey
Yizhang Jin, Jian Li, Yexin Liu, Tianjun Gu, Kai Wu, Zhengkai Jiang, Muyang He, Bo Zhao, Xin Tan, Zhenye Gan, Yabiao Wang, Chengjie Wang, and Lizhuang Ma · 2024
Closest in time.
Droid: A large-scale in-the-wild robot manipulation dataset
Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al · 2024
Closest in time.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, et al · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, Zhaoyang Zeng, Hao Zhang, Feng Li, Jie Yang, Hongyang Li, Qing Jiang, and Lei Zhang · 2024
Closest in time.
Octo: An open-source generalist robot policy
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, et al · 2024
Closest in time.
Manifoundation model for general-purpose robotic manipulation of contact synthesis with arbitrary objects and robots
Zhixuan Xu, Chongkai Gao, Zixuan Liu, Gang Yang, Chenrui Tie, Haozhuo Zheng, Haoyu Zhou, Weikun Peng, Debang Wang, Tianyi Chen, Zhouliang Yu, and Lin Shao · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al · 2024
Closest in time.
Jianhao Yuan, Shuyang Sun, Daniel Omeiza, Bo Zhao, Paul Newman, Lars Kunze, and Matthew Gadd · 2024
Closest in time.
Wordepth: Variational language prior for monocular depth estimation
Ziyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park, Yangchao Wu, Stefano Soatto, Byung-Woo Hong, Dong Lao, and Alex Wong · 2024
Closest in time.