Fetching the paper…
Reading the bibliography…
Spatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners.
On the theory of filter amplifiers
Stephen Butterworth et al · 1930
Earlier work this paper cites.
Aspects of the Theory of Syntax
Noam Chomsky · 1965
Earlier work this paper cites.
Up/down, front/back, left/right. a contrastive study of hausa and english
Clifford Hill · 1982
Earlier work this paper cites.
Parsing surrounding space into regions
Nancy Franklin, Linda A Henkel, and Thomas Zangas · 1995
Earlier work this paper cites.
Spatial language and spatial representation
William G Hayward and Michael J Tarr · 1995
Earlier work this paper cites.
Frames of reference and molyneux’s question: Crosslinguistic evidence
Stephen C Levinson · 1996
Earlier work this paper cites.
A computational analysis of the apprehension of spatial relations
Gordon D Logan and Daniel D Sadler · 1996
Earlier work this paper cites.
The influence of reference frame selection on spatial template construction
Laura A Carlson-Radvansky and Gordon D Logan · 1997
Earlier work this paper cites.
Formal models for cognition—taxonomy of spatial location description and frames of reference
Andrew U Frank · 1998
Earlier work this paper cites.
Spatial language and spatial representation: A cross-linguistic comparison
Edward Munnich, Barbara Landau, and Barbara Anne Dosher · 2001
Earlier work this paper cites.
Grounding spatial language in perception: an empirical and computational investigation
Terry Regier and Laura A Carlson · 2001
Earlier work this paper cites.
Space in language and cognition: Explorations in cognitive diversity , volume 5
Stephen C Levinson · 2003
Earlier work this paper cites.
Contextual, functional, and geometric components in the semantics of projective terms
Carola Eschenbach · 2004
Earlier work this paper cites.
Can language restructure cognition? the case for space
Asifa Majid, Melissa Bowerman, Sotaro Kita, Daniel BM Haun, and Stephen C Levinson · 2004
Earlier work this paper cites.
Identifying objects on the basis of spatial contrast: An empirical study
Thora Tenbrink · 2004
Earlier work this paper cites.
Spatial reference in linguistic human-robot interaction: Iterative, empirically supported development of a model of projective relations
Reinhard Moratz and Thora Tenbrink · 2006
Earlier work this paper cites.
Deixis, gesture, and cognition in spatial frame of reference typology
Eve Danziger · 2010
Earlier work this paper cites.
Spacing and orientation in co-present interaction
Adam Kendon · 2010
Earlier work this paper cites.
Ambiguities in spatial language understanding in situated human robot dialogue
Changsong Liu, Jacob Walker, and Joyce Y Chai · 2010
Earlier work this paper cites.
Evidence from an emerging sign language reveals that language supports spatial cognition
Jennie E Pyers, Anna Shusterman, Ann Senghas, Elizabeth S Spelke, and Karen Emmorey · 2010
Earlier work this paper cites.
Spatial frames of reference in mesoamerican languages
Carolyn O’Meara and Gabriela Pérez Báez · 2011
Earlier work this paper cites.
Children’s spatial thinking: Does talk about the spatial world matter?
Shannon M Pruden, Susan C Levine, and Janellen Huttenlocher · 2011
Earlier work this paper cites.
Psychology of spatial cognition
Luca Tommasi and Bruno Laeng · 2012
Earlier work this paper cites.
Development of spatial cognition
Marina Vasilyeva and Stella F Lourenco · 2012
Earlier work this paper cites.
Spatial language facilitates spatial cognition: Evidence from children who lack language input
Dedre Gentner, Asli Özyürek, Özge Gürcanli, and Susan Goldin-Meadow · 2013
Earlier work this paper cites.
The cultural transmission of spatial cognition: Evidence from a large-scale study
Jürgen Bohnemeyer, Katharine Donelson, Randi Tucker, Elena Benedicto, Alejandra Capistrán Garza, Alyson Eggleston, Néstor Hernández Green, María de Jesús Selene Hernández Gómez, Samuel Herrera Castro, Carolyn O’Meara, et al · 2014
Cited alongside, same era.
Blender - a 3d modelling and rendering package
Blender Online Community · 2016
Cited alongside, same era.
Frames of reference in spatial language acquisition
Anna Shusterman and Peggy Li · 2016
Cited alongside, same era.
Language to action: towards interactive task learning with physical agents
Joyce Y Chai, Qiaozi Gao, Lanbo She, Shaohua Yang, Sari Saba-Sadiya, and Guangyue Xu · 2018
Cited alongside, same era.
Being In Front
Andrea Bender, Sarah Teige-Mocigemba, Annelie Rothe-Wulf, Miriam Seel, and Sieghard Beller · 2020
Cited alongside, same era.
SPARTQA: A textual question answering benchmark for spatial reasoning
mblip: Efficient bootstrapping of multilingual vision-llms
Gregor Geigle, Abhay Jain, Radu Timofte, and Goran Glavaš · 2024
Closest in time.
Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, et al · 2024
Closest in time.
Ego3dt: Tracking every 3d object in ego-centric videos
Shengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun, Wendi Hu, Jieyang Zhou, Yixian Zhao, Qi Li, Yizhou Wang, Xi Li, et al · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al · 2024
Closest in time.
MMTom-QA: Multimodal theory of mind question answering
Chuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua Tenenbaum, and Tianmin Shu · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roshanak Mirzaee, Hossein Rajaby Faghihi, Qiang Ning, and Parisa Kordjamshidi · 2021
Cited alongside, same era.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
Transfer learning with synthetic corpora for spatial role labeling and reasoning
Roshanak Mirzaee and Parisa Kordjamshidi · 2022
Cited alongside, same era.
What artificial neural networks can tell us about human language acquisition
Alex Warstadt and Samuel R Bowman · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Closest in time.
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia · 2024
Closest in time.
Multilingual diversity improves vision-language representations
Thao Nguyen, Matthew Wallingford, Sebastin Santy, Wei-Chiu Ma, Sewoong Oh, Ludwig Schmidt, Pang Wei Koh, and Ranjay Krishna · 2024
Closest in time.
Spatial frames of reference in dholuo
Awino Ogelo and Emanuel Bylund · 2024
Closest in time.
Hello gpt-4o, May 2024
OpenAI · 2024
Closest in time.
Grounding multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, Qixiang Ye, and Furu Wei · 2024
Closest in time.
Affordancellm: Grounding affordance from vision language models
Shengyi Qian, Weifeng Chen, Min Bai, Xiong Zhou, Zhuowen Tu, and Li Erran Li · 2024
Closest in time.
Towards grounded visual spatial reasoning in multi-modal vision language models
Navid Rajabi and Jana Kosecka · 2024
Closest in time.
Glamm: Pixel grounding large multimodal model
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S Khan · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Cvqa: Culturally-diverse multilingual visual question answering benchmark
David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo, Teresa Lynn, Injy Hamed, Aditya Nanda Kishore, Aishik Mandal, Alina Dragonetti, Artem Abzaliev, Atnafu Lambebo Tonja, et al · 2024
Closest in time.
Benchmarks as microscopes: A call for model metrology
Michael Saxon, Ari Holtzman, Peter West, William Yang Wang, and Naomi Saphra · 2024
Closest in time.
Learning Language Structures through Grounding
Haoyue Freda Shi · 2024
Closest in time.
Gsva: Generalized segmentation via multimodal large language models
Zhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan, Shiji Song, and Gao Huang · 2024
Closest in time.
3d-grand: A million-scale dataset for 3d-llms with better grounding and less hallucination
Jianing Yang, Xuweiyi Chen, Nikhil Madaan, Madhavan Iyengar, Shengyi Qian, David F Fouhey, and Joyce Chai · 2024
Closest in time.
Robopoint: A vision-language model for spatial affordance prediction in robotics
Wentao Yuan, Jiafei Duan, Valts Blukis, Wilbert Pumacay, Ranjay Krishna, Adithyavairavan Murali, Arsalan Mousavian, and Dieter Fox · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Closest in time.
Groundhog: Grounding large language models to holistic segmentation
Yichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah, Qiaozi Gao, and Joyce Chai · 2024
Closest in time.
Neuro-symbolic training for reasoning over spatial language
Tanawan Premsri and Parisa Kordjamshidi · 2025
Closest in time.
Logical forms complement probability in understanding language model (and human) performance
Yixuan Wang and Freda Shi · 2025
Closest in time.