Fetching the paper…
Reading the bibliography…
Over the past year, the development of large language models (LLMs) has brought spatial intelligence into focus, with much attention on vision-based embodied intelligence.
Cognitive maps in rats and men
Edward C Tolman · 1948
Earlier work this paper cites.
Memory, amnesia and the hippocampal system
NJ Cohen · 1993
Earlier work this paper cites.
Place cells, grid cells, and the brain’s spatial representation system
Edvard I Moser, Emilio Kropff, and May-Britt Moser · 2008
Earlier work this paper cites.
Can we reconcile the declarative memory and spatial navigation views on hippocampal function?
Howard Eichenbaum and Neal J Cohen · 2014
Earlier work this paper cites.
The cognitive map in humans: spatial navigation and beyond
Russell A Epstein, Eva Zita Patai, Joshua B Julian, and Hugo J Spiers · 2017
Earlier work this paper cites.
Neurobiology of schemas and schema-mediated memory
Asaf Gilboa and Hannah Marlatte · 2017
Earlier work this paper cites.
Spatial representation in the hippocampal formation: a history
Edvard I Moser, May-Britt Moser, and Bruce L McNaughton · 2017
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel · 2019
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Earlier work this paper cites.
The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation
James CR Whittington, Timothy H Muller, Shirley Mark, Guifen Chen, Caswell Barry, Neil Burgess, and Timothy EJ Behrens · 2020
Earlier work this paper cites.
Embodied intelligence via learning and evolution
Agrim Gupta, Silvio Savarese, et al · 2021
Earlier work this paper cites.
Spatial thinking, cognitive mapping, and spatial awareness
Toru Ishikawa · 2021
Earlier work this paper cites.
Geographic question answering: challenges, uniqueness, classification, and future directions
Gengchen Mai, Krzysztof Janowicz, Rui Zhu, Ling Cai, and Ni Lao · 2021
Earlier work this paper cites.
Skilful precipitation nowcasting using deep generative models of radar
Suman Ravuri, Karel Lenc, Matthew Willson, Dmitry Kangin, Remi Lam, Piotr Mirowski, Megan Fitzsimons, Maria Athanassiadou, Sheleem Kashem, Sam Madge, et al · 2021
Earlier work this paper cites.
Relating transformers to models and neural representations of the hippocampal formation
James CR Whittington, Joseph Warren, and Timothy EJ Behrens · 2021
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan · 2022
Earlier work this paper cites.
Factuality enhanced language models for open-ended text generation
Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale N Fung, Mohammad Shoeybi, and Bryan Catanzaro · 2022
Earlier work this paper cites.
Developing knowledge graph based system for urban computing
Yu Liu, Jingtao Ding, and Yong Li · 2022
Earlier work this paper cites.
Are large language models geospatially knowledgeable?
Prabin Bhandari, Antonios Anastasopoulos, and Dieter Pfoser · 2023
Earlier work this paper cites.
Accurate medium-range global weather forecasting with 3d neural networks
Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian · 2023
Earlier work this paper cites.
Fuxi: A cascade machine learning forecasting system for 15-day global weather forecast
Lei Chen, Xiaohui Zhong, et al · 2023
Earlier work this paper cites.
From cognitive maps to spatial schemas
Delaram Farzanfar, Hugo J Spiers, Morris Moscovitch, and R Shayna Rosenbaum · 2023
Earlier work this paper cites.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al · 2023
Earlier work this paper cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei · 2023
Earlier work this paper cites.
Geomverse: A systematic evaluation of large models for geometric reasoning
Mehran Kazemi, Hamidreza Alvari, Ankit Anand, Jialin Wu, Xi Chen, and Radu Soricut · 2023
Earlier work this paper cites.
Large language models as traffic signal control agents: Capacity and opportunity
Siqi Lai, Zhao Xu, Weijia Zhang, Hao Liu, and Hui Xiong · 2023
Earlier work this paper cites.
Urbankg: An urban knowledge graph system
Yu Liu, Jingtao Ding, Yanjie Fu, and Yong Li · 2023
Earlier work this paper cites.
Sphere2Vec: A general-purpose location representation learning over a spherical surface for large-scale geospatial predictions
Gengchen Mai, Yao Xuan, et al · 2023
Earlier work this paper cites.
Geollm: Extracting geospatial knowledge from large language models
Rohin Manvi, Samar Khanna, Gengchen Mai, Marshall Burke, David Lobell, and Stefano Ermon · 2023
Earlier work this paper cites.
Exploring and improving the spatial reasoning abilities of large language models
Manasi Sharma · 2023
Earlier work this paper cites.
Generating explanations for embodied action decision from visual observation
Xiaohan Wang, Yuehu Liu, Xinhang Song, Beibei Wang, and Shuqiang Jiang · 2023
Earlier work this paper cites.
Evaluating spatial understanding of large language models
Yutaro Yamada, Yihan Bao, Andrew K Lampinen, Jungo Kasai, and Ilker Yildirim · 2023
Cited alongside, same era.
Geogpt: understanding and processing geospatial tasks through an autonomous gpt
Yifan Zhang, Cheng Wei, Shangyou Wu, Zhengting He, and Wenhao Yu · 2023
Cited alongside, same era.
Skilful nowcasting of extreme precipitation with nowcastnet
Yuchen Zhang, Mingsheng Long, Kaiyuan Chen, Lanxiang Xing, Ronghua Jin, Michael I Jordan, and Jianmin Wang · 2023
Cited alongside, same era.
How do large language models capture the ever-changing world knowledge? a review of recent advances
Zihan Zhang, Meng Fang, Ling Chen, Mohammad-Reza Namazi-Rad, and Jun Wang · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Mp5: A multi-modal open-ended embodied system in minecraft via active perception
Yiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin, Lu Sheng, Ruimao Zhang, Yu Qiao, and Jing Shao · 2024
Later among the works it cites.
Charting new territories: Exploring the geographic and geospatial capabilities of multimodal llms
Jonathan Roberts, Timo Lüddecke, Rehan Sheikh, Kai Han, and Samuel Albanie · 2024
Later among the works it cites.
Velma: Verbalization embodiment of llm agents for vision and language navigation in street view
Raphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu, Stefan Riezler, and William Yang Wang · 2024
Later among the works it cites.
Beyond imitation: Generating human mobility from context-aware reasoning with large language models
Chenyang Shao, Fengli Xu, Bingbing Fan, Jingtao Ding, Yuan Yuan, Meng Wang, and Yong Li · 2024
Later among the works it cites.
Llmdiff: Diffusion model using frozen llm transformers for precipitation nowcasting
Lei She, Chenghong Zhang, Xin Man, and Jie Shao · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al · 2023
Cited alongside, same era.
Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun · 2024
Cited alongside, same era.
Spatialbot: Precise spatial understanding with vision language models
Wenxiao Cai, Yaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li, Wankou Yang, Hao Dong, and Bo Zhao · 2024
Cited alongside, same era.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Boyuan Chen, Zhuo Xu, et al · 2024
Cited alongside, same era.
Understanding world or predicting future? a comprehensive survey of world models
Jingtao Ding, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, Hongyuan Su, Nian Li, Nicholas Sukiennik, et al · 2024
Cited alongside, same era.
Citygpt: Empowering urban spatial cognition of large language models
Jie Feng, Yuwei Du, Tianhui Liu, Siqi Guo, Yuming Lin, and Yong Li · 2024
Cited alongside, same era.
Agentmove: Predicting human mobility anywhere using large language model based agentic framework
Jie Feng, Yuwei Du, Jie Zhao, and Yong Li · 2024
Cited alongside, same era.
Citybench: Evaluating the capabilities of large language models for urban tasks, 2024
Jie Feng, Jun Zhang, Tianhui Liu, Xin Zhang, Tianjian Ouyang, Junbo Yan, Yuwei Du, Siqi Guo, and Yong Li · 2024
Cited alongside, same era.
Later among the works it cites.
Sangmim Song, Sarath Kodagoda, Amal Gunatilake, Marc G Carmichael, Karthick Thiyagarajan, and Jodi Martin · 2024
Later among the works it cites.
Large language models as urban residents: An llm agent framework for personal mobility generation
Jiawei Wang, Renhe Jiang, Chuang Yang, Zengqing Wu, Makoto Onizuka, Ryosuke Shibasaki, Noboru Koshizuka, and Chuan Xiao · 2024
Later among the works it cites.
Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai
Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen, Wenwei Zhang, Kai Chen, Tianfan Xue, et al · 2024
Later among the works it cites.
Torchspatial: A location encoding framework and benchmark for spatial representation learning
Nemin Wu, Qian Cao, et al · 2024
Later among the works it cites.
Refound: Crafting a foundation model for urban region understanding upon language and visual foundations
Congxi Xiao, Jingbo Zhou, Yixiong Xiao, Jizhou Huang, and Hui Xiong · 2024
Later among the works it cites.
Flame: Learning to navigate with multimodal llm in urban environments
Yunzhe Xu, Yiyuan Pan, Zhe Liu, and Hesheng Wang · 2024
Later among the works it cites.
Geopredict-llm: Intelligent tunnel advanced geological prediction by reprogramming large language models
Zhenhao Xu, Zhaoyang Wang, Shucai Li, Xiao Zhang, and Peng Lin · 2024
Later among the works it cites.
Georeasoner: Reasoning on geospatially grounded context for natural language understanding
Yibo Yan and Joey Lee · 2024
Later among the works it cites.
Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web
Yibo Yan, Haomin Wen, Siru Zhong, Wei Chen, Haodong Chen, Qingsong Wen, Roger Zimmermann, and Yuxuan Liang · 2024
Later among the works it cites.
Llmi3d: Empowering llm with 3d perception from a single 2d image
Fan Yang, Sicheng Zhao, Yanhao Zhang, Haoxiang Chen, Hui Chen, Wenbo Tang, Haonan Lu, Pengfei Xu, Zhenyu Yang, Jungong Han, et al · 2024
Later among the works it cites.
Thinking in space: How multimodal large language models see, remember, and recall spaces, 2024
Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie · 2024
Later among the works it cites.
Ruochu Yang, Fumin Zhang, and Mengxue Hou · 2024
Later among the works it cites.
3d-mem: 3d scene memory for embodied exploration and reasoning
Yuncong Yang, Han Yang, Jiachen Zhou, Peihao Chen, Hongxin Zhang, Yilun Du, and Chuang Gan · 2024
Later among the works it cites.
Mineagent: Towards remote-sensing mineral exploration with multimodal large language models
Beibei Yu, Tao Shen, Hongbin Na, Ling Chen, and Denqi Li · 2024
Later among the works it cites.
Rag-guided large language models for visual spatial description with adaptive hallucination corrector
Jun Yu, Yunxiang Zhang, Zerui Zhang, Zhao Yang, Gongpeng Zhao, Fengzhao Sun, Fanrui Zhang, Qingsong Liu, Jianqing Sun, Jiaen Liang, et al · 2024
Later among the works it cites.
Qingbin Zeng, Qinglong Yang, Shunan Dong, Heming Du, Liang Zheng, Fengli Xu, and Yong Li · 2024
Later among the works it cites.
Geoeval: benchmark for evaluating llms and multi-modal models on geometry problem-solving
Jiaxin Zhang, Zhongzhi Li, Mingliang Zhang, Fei Yin, Chenglin Liu, and Yashar Moshfeghi · 2024
Later among the works it cites.
Artificial intelligence for geoscience: Progress, challenges and perspectives
Tianjie Zhao, Sheng Wang, et al · 2024
Later among the works it cites.
Video-3d llm: Learning position-aware video representation for 3d scene understanding
Duo Zheng, Shijia Huang, and Liwei Wang · 2024
Later among the works it cites.
Topv-nav: Unlocking the top-view spatial reasoning potential of mllm for zero-shot object navigation
Linqing Zhong, Chen Gao, Zihan Ding, Yue Liao, and Si Liu · 2024
Later among the works it cites.
Navgpt: Explicit reasoning in vision-and-language navigation with large language models
Gengze Zhou, Yicong Hong, and Qi Wu · 2024
Later among the works it cites.
Large language model for participatory urban planning
Zhilun Zhou, Yuming Lin, Depeng Jin, and Yong Li · 2024
Later among the works it cites.
Allocentric and egocentric spatial representations coexist in rodent medial entorhinal cortex
Xiaoyang Long, Daniel Bush, Bin Deng, Neil Burgess, and Sheng-Jia Zhang · 2025
Closest in time.
Wenrui Xu, Dalin Lyu, Weihang Wang, Jie Feng, Chen Gao, and Yong Li · 2025
Closest in time.
Navgpt-2: Unleashing navigational reasoning capability for large vision-language models
Gengze Zhou, Yicong Hong, Zun Wang, Xin Eric Wang, and Qi Wu · 2025
Closest in time.