Fetching the paper…
Reading the bibliography…
Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored.
Challenging the egocentric view of coordinated perceiving, acting, and knowing
Michael J Richardson, Kerry L Marsh, and RC Schmidt. 2010 · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David Chen and William B Dolan. 2011 · 2011
Earlier work this paper cites.
Multi-target detection and tracking from a single camera in unmanned aerial vehicles (uavs)
Jing Li, Dong Hye Ye, Timothy Chung, Mathias Kolsch, Juan Wachs, and Charles Bouman. 2016 · 2016
Earlier work this paper cites.
Mobile three-dimensional maps for wayfinding in large and complex buildings: Empirical comparison of first-person versus third-person perspective
Stefano Burigat, Luca Chittaro, and Riccardo Sioni. 2017 · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017 · 2017
Earlier work this paper cites.
Embodied cognition
Lawrence Shapiro. 2019 · 2019
Earlier work this paper cites.
Object goal navigation using goal-oriented semantic exploration
Devendra Singh Chaplot, Dhiraj Prakashchand Gandhi, Abhinav Gupta, and Russ R Salakhutdinov. 2020 · 2020
Earlier work this paper cites.
Vision-language navigation with self-supervised auxiliary reasoning tasks
Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang. 2020 · 2020
Earlier work this paper cites.
Embodied intelligence via learning and evolution
Agrim Gupta, Silvio Savarese, Surya Ganguli, and Li Fei-Fei. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Human-centric spatio-temporal video grounding with visual transformers
Zongheng Tang, Yue Liao, Si Liu, Guanbin Li, Xiaojie Jin, Hongxu Jiang, Qian Yu, and Dong Xu. 2021 · 2021
Earlier work this paper cites.
Detection, tracking, and counting meets drones in crowds: A benchmark
Longyin Wen, Dawei Du, Pengfei Zhu, Qinghua Hu, Qilong Wang, Liefeng Bo, and Siwei Lyu. 2021 · 2021
Earlier work this paper cites.
Detection and tracking meet drones challenge
Pengfei Zhu, Longyin Wen, Dawei Du, Xiao Bian, Heng Fan, Qinghua Hu, and Haibin Ling. 2021 · 2021
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar. 2022 · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022 · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta. 2022 · 2022
Earlier work this paper cites.
An embodied generalist agent in 3d world
Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. 2023 · 2023
Earlier work this paper cites.
Large multimodal models: Notes on cvpr 2023 tutorial
Chunyuan Li. 2023 · 2023
Earlier work this paper cites.
Aerialvln: Vision-and-language navigation for uavs
Shubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang, Yanning Zhang, and Qi Wu. 2023 · 2023
Earlier work this paper cites.
Where are we in the search for an artificial visual cortex for embodied intelligence?
Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Tingfan Wu, Jay Vakil, et al. 2023 · 2023
Cited alongside, same era.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. 2023 · 2023
Cited alongside, same era.
Building a cognitive science of human variation: Individual differences in spatial navigation
Nora S Newcombe, Mary Hegarty, and David Uttal. 2023 · 2023
Cited alongside, same era.
Scene graph contrastive learning for embodied navigation
Kunal Pratap Singh, Jordi Salvador, Luca Weihs, and Aniruddha Kembhavi. 2023 · 2023
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su. 2023 · 2023
Cited alongside, same era.
Videoinsta: Zero-shot long video understanding via informative spatial-temporal reasoning with llms
Ruotong Liao, Max Erler, Huiyu Wang, Guangyao Zhai, Gengyuan Zhang, Yunpu Ma, and Volker Tresp. 2024 · 2024
Later among the works it cites.
Lingoqa: Visual question answering for autonomous driving
Ana-Maria Marcu, Long Chen, Jan Hünermann, Alice Karnsund, Benoit Hanotte, Prajwal Chidananda, Saurabh Nair, Vijay Badrinarayanan, Alex Kendall, Jamie Shotton, et al. 2024 · 2024
Later among the works it cites.
Does spatial cognition emerge in frontier models?
Santhosh Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, and Vladlen Koltun. 2024 · 2024
Later among the works it cites.
Exploring efficient foundational multi-modal models for video summarization
Karan Samel, Apoorva Beedu, Nitish Sontakke, and Irfan Essa. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yunlong Tang, Jing Bi, Siting Xu, Luchuan Song, Susan Liang, Teng Wang, Daoan Zhang, Jie An, Jingyang Lin, Rongyi Zhu, et al. 2023 · 2023
Cited alongside, same era.
Urban generative intelligence (ugi): A foundational platform for agents in embodied city environment
Fengli Xu, Jun Zhang, Chen Gao, Jie Feng, and Yong Li. 2023 · 2023
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2023 · 2023
Cited alongside, same era.
Hourvideo: 1-hour video-language understanding
Keshigeyan Chandrasegaran, Agrim Gupta, Lea M Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei. 2024 · 2024
Cited alongside, same era.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. 2024 · 2024
Cited alongside, same era.
Videgothink: Assessing egocentric video understanding capabilities for embodied ai
Sijie Cheng, Kechen Fang, Yangyang Yu, Sicheng Zhou, Bohao Li, Ye Tian, Tingguang Li, Lei Han, and Yang Liu. 2024 · 2024
Cited alongside, same era.
Understanding world or predicting future? a comprehensive survey of world models
Jingtao Ding, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, Hongyuan Su, Nian Li, Nicholas Sukiennik, et al. 2024 · 2024
Cited alongside, same era.
Haochen Shi, Zhiyuan Sun, Xingdi Yuan, Marc-Alexandre Côté, and Bang Liu. 2024 · 2024
Later among the works it cites.
Drivelm: Driving with graph visual question answering
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Jens Beißwenger, Ping Luo, Andreas Geiger, and Hongyang Li. 2024 · 2024
Later among the works it cites.
Moviechat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. 2024 · 2024
Later among the works it cites.
Alanavlm: A multimodal embodied ai foundation model for egocentric video understanding
Alessandro Suglia, Claudio Greco, Katie Baker, Jose L Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon. 2024 · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024 · 2024
Later among the works it cites.
Thinking in space: How multimodal large language models see, remember, and recall spaces
Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. 2024 · 2024
Later among the works it cites.
Fanglong Yao, Yuanchang Yue, Youzhi Liu, Xian Sun, and Kun Fu. 2024 · 2024
Later among the works it cites.
Qingbin Zeng, Qinglong Yang, Shunan Dong, Heming Du, Liang Zheng, Fengli Xu, and Yong Li. 2024 · 2024
Later among the works it cites.
Navid: Video-based vlm plans the next step for vision-and-language navigation
Jiazhao Zhang, Kunyu Wang, Rongtao Xu, Gengze Zhou, Yicong Hong, Xiaomeng Fang, Qi Wu, Zhizheng Zhang, and He Wang. 2024 · 2024
Later among the works it cites.
Navgpt: Explicit reasoning in vision-and-language navigation with large language models
Gengze Zhou, Yicong Hong, and Qi Wu. 2024 · 2024
Later among the works it cites.
Qwen documentation
Alibaba Cloud. 2025 · 2025
Closest in time.
Gemini api documentation
Google. 2025 · 2025
Closest in time.
Openai api documentation
OpenAI. 2025 · 2025
Closest in time.
Wenrui Xu, Dalin Lyu, Weihang Wang, Jie Feng, Chen Gao, and Yong Li. 2025 · 2025
Closest in time.