Fetching the paper…
Reading the bibliography…
Cross-view correspondence is a fundamental capability for spatial understanding and embodied AI.
Mental rotation of three-dimensional objects
Roger N Shepard and Jacqueline Metzler · 1971
Earlier work this paper cites.
Human spatial representation: Insights from animals
Ranxiao Frances Wang and Elizabeth S Spelke · 2002
Earlier work this paper cites.
A diagram is worth a dozen images
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarjan, Meet Shah, Yu Jiang, Xinlei Chen, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
University-1652: A multi-view multi-source benchmark for drone-based geo-localization
Zhedong Zheng, Yunchao Wei, and Yi Yang · 2020
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen · 2023
Earlier work this paper cites.
Scannet++: A high-fidelity dataset of 3d indoor scenes
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai · 2023
Earlier work this paper cites.
Benchlmm: Benchmarking cross-style visual capability of large multimodal models
Rizhao Cai, Zirui Song, Dayan Guan, Zhenhao Chen, Yaohang Li, Xing Luo, Chenyu Yi, and Alex Kot · 2024
Earlier work this paper cites.
Mvinpainter: learning multi-view consistent inpainting to bridge 2d and 3d editing
Chenjie Cao, Chaohui Yu, Fan Wang, Xiangyang Xue, and Yanwei Fu · 2024
Earlier work this paper cites.
Spatialrgpt: Grounded spatial reasoning in vision-language models
An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu · 2024
Earlier work this paper cites.
Large spatial model: End-to-end unposed images to semantic 3d
Zhiwen Fan, Jian Zhang, Wenyan Cong, Peihao Wang, Renjie Li, Kairun Wen, Shijie Zhou, Achuta Kadambi, Zhangyang Wang, Danfei Xu, et al · 2024
Earlier work this paper cites.
Blink: Multimodal large language models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna · 2024
Earlier work this paper cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al · 2024
Earlier work this paper cites.
Improved visual grounding through self-consistent explanations
Ruozhen He, Paola Cascante-Bonilla, Ziyan Yang, Alexander C Berg, and Vicente Ordonez · 2024
Earlier work this paper cites.
Egosim: An egocentric multi-view simulator and real dataset for body-worn cameras during motion and activity
Dominik Hollidt, Paul Streli, Jiaxi Jiang, Yasaman Haghighi, Changlin Qian, Xintong Liu, and Christian Holz · 2024
Cited alongside, same era.
Multi-agent collaborative perception via motion-aware robust communication network
Shixin Hong, Yu Liu, Zhi Li, Shaohui Li, and You He · 2024
Cited alongside, same era.
Towards robust estimation of human intention hierarchy in robot teleoperation
Nikki Lijing Kuang, Songpo Li, and Soshi Iba · 2024
Cited alongside, same era.
Rapid motor adaptation for robotic manipulator arms
Yichao Liang, Kevin Ellis, and Joao Henriques · 2024
Cited alongside, same era.
Real-time simulated avatar from head-mounted sensors
Zhengyi Luo, Jinkun Cao, Rawal Khirodkar, Alexander Winkler, Kris Kitani, and Weipeng Xu · 2024
Cited alongside, same era.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Embodied ai for smart robotic cells in manufacturing applications
Satyandra K Gupta · 2025
Closest in time.
Your large vision-language model only needs a few attention heads for visual grounding
Seil Kang, Jinyeong Kim, Junhyeok Kim, and Seong Jae Hwang · 2025
Closest in time.
Multi-sensor object anomaly detection: Unifying appearance, geometry, and internal properties
Wenqiao Li, Bozhong Zheng, Xiaohao Xu, Jinye Gan, Fading Lu, Xiang Li, Na Ni, Zheng Tian, Xiaonan Huang, Shenghua Gao, et al · 2025
Closest in time.
Bingyi Liu, Jian Teng, Hongfei Xue, Enshu Wang, Chuanhui Zhu, Pu Wang, and Libing Wu · 2025
Closest in time.
Egoprompt: Prompt pool learning for egocentric action recognition
Huaihai Lyu, Chaofan Chen, Yuheng Ji, and Changsheng Xu · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Adithya Jairam Vedagiri IYER, Sai Charitha Akula, Shusheng Yang, Jihan Yang, Manoj Middepogu, Ziteng Wang, et al · 2024
Cited alongside, same era.
Surgicai: A hierarchical platform for fine-grained surgical policy learning and benchmarking
Jin Wu, Haoying Zhou, Peter Kazanzides, Adnan Munawar, and Anqi Liu · 2024
Cited alongside, same era.
V2x-vitv2: Improved vision transformers for vehicle-to-everything cooperative perception
Runsheng Xu, Chia-Ju Chen, Zhengzhong Tu, and Ming-Hsuan Yang · 2024
Cited alongside, same era.
Boosting weakly supervised referring image segmentation via progressive comprehension
Zaiquan Yang, Yuhao Liu, Jiaying Lin, Gerhard Hancke, and Rynson Lau · 2024
Cited alongside, same era.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Cited alongside, same era.
Introducing claude 4
Anthropic · 2025
Cited alongside, same era.
Met3r: Measuring multi-view consistency in generated images
Mohammad Asim, Christopher Wewer, Thomas Wimmer, Bernt Schiele, and Jan Eric Lenssen · 2025
Cited alongside, same era.
Closest in time.
Which viewpoint shows it best? language for weakly supervising view selection in multi-view instructional videos
Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan, and Kristen Grauman · 2025
Closest in time.
Openai o3 and o4-mini system card
OpenAI · 2025
Closest in time.
R-vlm: Region-aware vision language model for precise gui grounding
Joonhyung Park, Peng Tang, Sagnik Das, Srikar Appalaraju, Kunwar Yashraj Singh, R Manmatha, and Shabnam Ghadar · 2025
Closest in time.
SAM 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollar, and Christoph Feichtenhofer · 2025
Closest in time.
SAT: Dynamic spatial aptitude training for multimodal language models
Arijit Ray, Jiafei Duan, Ellis L Brown II, Reuben Tan, Dina Bashkirova, Rose Hendrix, Kiana Ehsani, Aniruddha Kembhavi, Bryan A. Plummer, Ranjay Krishna, Kuo-Hao Zeng, and Kate Saenko · 2025
Closest in time.
Hazards in daily life? enabling robots to proactively detect and resolve anomalies
Zirui Song, Guangxian Ouyang, Meng Fang, Hongbin Na, Zijing Shi, Zhenhao Chen, Fu Yujie, Zeyu Zhang, Shiyu Jiang, Miao Fang, et al · 2025
Closest in time.
Learning to detect objects from multi-agent lidar scans without manual labels
Qiming Xia, Wenkai Lin, Haoen Xiang, Xun Huang, Siheng Chen, Zhen Dong, Cheng Wang, and Chenglu Wen · 2025
Closest in time.
Mc-bench: A benchmark for multi-context visual grounding in the era of mllms
Yunqiu Xu, Linchao Zhu, and Yi Yang · 2025
Closest in time.
Seeing from another perspective: Evaluating multi-view understanding in mllms
Chun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta-Ying Cheng, Ruoyu Wang, Tianzhe Chu, Yuexiang Zhai, Yubei Chen, Shenghua Gao, and Yi Ma · 2025
Closest in time.
Robopoint: A vision-language model for spatial affordance prediction in robotics
Wentao Yuan, Jiafei Duan, Valts Blukis, Wilbert Pumacay, Ranjay Krishna, Adithyavairavan Murali, Arsalan Mousavian, and Dieter Fox · 2025
Closest in time.
Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark
Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, et al · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, et al · 2025
Closest in time.
Cooptrack: Exploring end-to-end learning for efficient cooperative sequential perception
Jiaru Zhong, Jiahao Wang, Jiahui Xu, Xiaofan Li, Zaiqing Nie, and Haibao Yu · 2025
Closest in time.