Fetching the paper…
Reading the bibliography…
Spatial Planning is a crucial part in the field of spatial intelligence, which requires the understanding and planning about object arrangements in space perspective.
Mental rotation of three-dimensional objects
Roger N Shepard and Jacqueline Metzler · 1971
Earlier work this paper cites.
Mental rotations, a group test of three-dimensional spatial visualization
Steven G Vandenberg and Allan R Kuse · 1978
Earlier work this paper cites.
Mental rotation: effects of dimensionality of objects and type of task
Shenna Shepard and Douglas Metzler · 1988
Earlier work this paper cites.
Spatial reasoning
Ruth MJ Byrne and Philip N Johnson-Laird · 1989
Earlier work this paper cites.
Geometry and spatial reasoning
Douglas H Clements and Michael T Battista · 1992
Earlier work this paper cites.
Creative innovation: possible brain mechanisms
Kenneth M Heilman, Stephen E Nadeau, and David O Beversdorf · 2003
Earlier work this paper cites.
Components of spatial intelligence
Mary Hegarty · 2010
Earlier work this paper cites.
Commonsense reasoning and commonsense knowledge in artificial intelligence
Ernest Davis and Gary Marcus · 2015
Earlier work this paper cites.
On the ranking of a swiss system chess team tournament
László Csató · 2017
Earlier work this paper cites.
Acquiring common sense spatial knowledge through implicit spatial templates
Guillem Collell, Luc Van Gool, and Marie-Francine Moens · 2018
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
William H Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar · 2022
Earlier work this paper cites.
Open-ended evolution for minecraft building generation
Matthew Barthet, Antonios Liapis, and Georgios N Yannakakis · 2022
Earlier work this paper cites.
Things not written in text: Exploring spatial commonsense from visual signals
Xiao Liu, Da Yin, Yansong Feng, and Dongyan Zhao · 2022
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Earlier work this paper cites.
Gpt-4v(ision) system card, 2023
OpenAI · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Earlier work this paper cites.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Earlier work this paper cites.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen · 2023
Earlier work this paper cites.
Instruction-following evaluation for large language models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou · 2023
Cited alongside, same era.
OpenAI:Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, FlorenciaLeoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, HyungWon Chung, Dave Cummings, and Jeremiah Currier · 2023
Cited alongside, same era.
Artificial general intelligence: Humanity’s downturn or unlimited prosperity
Paolo Faraboschi, Eitan Frachtenberg, Phil Laplante, Dejan Milojicic, and Roberto Saracco · 2023
Cited alongside, same era.
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi · 2023
Mengfei Du, Binhao Wu, Zejun Li, Xuanjing Huang, and Zhongyu Wei · 2024
Later among the works it cites.
Shiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu, Yuan Li, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu-Gang Jiang, et al · 2024
Later among the works it cites.
Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks
Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, and Liqiang Nie · 2024
Later among the works it cites.
Teamcraft: A benchmark for multi-modal multi-agent systems in minecraft
Qian Long, Zhi Li, Ran Gong, Ying Nian Wu, Demetri Terzopoulos, and Xiaofeng Gao · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Steve-1: A generative model for text-to-behavior in minecraft
Shalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba, and Sheila McIlraith · 2023
Cited alongside, same era.
Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Ma, and Yitao Liang · 2023
Cited alongside, same era.
Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics
Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu Su, and Stan Birchfield · 2024
Cited alongside, same era.
Sat: Spatial aptitude training for multimodal language models
Arijit Ray, Jiafei Duan, Reuben Tan, Dina Bashkirova, Rose Hendrix, Kiana Ehsani, Aniruddha Kembhavi, Bryan A Plummer, Ranjay Krishna, Kuo-Hao Zeng, et al · 2024
Cited alongside, same era.
Seeground: See and ground for zero-shot open-vocabulary 3d visual grounding
Rong Li, Shijie Li, Lingdong Kong, Xulei Yang, and Junwei Liang · 2024
Cited alongside, same era.
Spartun3d: Situated spatial understanding of 3d world in large language models
Yue Zhang, Zhiyang Xu, Ying Shen, Parisa Kordjamshidi, and Lifu Huang · 2024
Cited alongside, same era.
Thinking in space: How multimodal large language models see, remember, and recall spaces
Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie · 2024
Cited alongside, same era.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia · 2024
Cited alongside, same era.
Shaofei Cai, Zihao Wang, Kewei Lian, Zhancun Mu, Xiaojian Ma, Anji Liu, and Yitao Liang · 2024
Later among the works it cites.
Minsstudio: A streamlined package for minecraft ai agent development
Shaofei Cai, Zhancun Mu, Kaichen He, Bowei Zhang, Xinyue Zheng, Anji Liu, and Yitao Liang · 2024
Later among the works it cites.
Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models
Zihao Wang, Shaofei Cai, Anji Liu, Yonggang Jin, Jinbing Hou, Bowei Zhang, Haowei Lin, Zhaofeng He, Zilong Zheng, Yaodong Yang, et al · 2024
Later among the works it cites.
Groot-2: Weakly supervised multi-modal instruction following agents
Shaofei Cai, Bowei Zhang, Zihao Wang, Haowei Lin, Xiaojian Ma, Anji Liu, and Yitao Liang · 2024
Later among the works it cites.
Mp5: A multi-modal open-ended embodied system in minecraft via active perception
Yiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin, Lu Sheng, Ruimao Zhang, Yu Qiao, and Jing Shao · 2024
Later among the works it cites.
Lego-puzzles: How good are mllms at multi-step spatial reasoning?
Kexian Tang, Junyao Gao, Yanhong Zeng, Haodong Duan, Yanan Sun, Zhening Xing, Wenran Liu, Kaifeng Lyu, and Kai Chen · 2025
Closest in time.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin · 2025
Closest in time.
Introducing gpt-4.1 in the api
OpenAI · 2025
Closest in time.
Sti-bench: Are mllms ready for precise spatial-temporal world understanding?
Yun Li, Yiming Zhang, Tao Lin, XiangRui Liu, Wenxiao Cai, Zheng Liu, and Bo Zhao · 2025
Closest in time.
Yuecheng Liu, Dafeng Chi, Shiguang Wu, Zhanguang Zhang, Yaochen Hu, Lingfeng Zhang, Yingxue Zhang, Shuang Wu, Tongtong Cao, Guowei Huang, et al · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model
Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Yan Ding, Zhigang Wang, JiaYuan Gu, Bin Zhao, Dong Wang, et al · 2025
Closest in time.
Mineworld: a real-time and open-source interactive world model on minecraft
Junliang Guo, Yang Ye, Tianyu He, Haoyu Wu, Yushu Jiang, Tim Pearce, and Jiang Bian · 2025
Closest in time.
Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy
Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, and Liqiang Nie · 2025
Closest in time.
Rocket-2: Steering visuomotor policy via cross-view goal alignment
Shaofei Cai, Zhancun Mu, Anji Liu, and Yitao Liang · 2025
Closest in time.
Cognitive traits shape the brain activity associated with mental rotation
Nadia M Bersier, Eleonora Fornari, Raffaella I Rumiati, and Silvio Ionta · 2025
Closest in time.
Yunfan Jiang, Ruohan Zhang, Josiah Wong, Chen Wang, Yanjie Ze, Hang Yin, Cem Gokmen, Shuran Song, Jiajun Wu, and Li Fei-Fei · 2025
Closest in time.