Fetching the paper…
Reading the bibliography…
We investigate the abilities of a representative set of Large language Models (LLMs) to reason about cardinal directions (CDs).
Guugu Yimithirr cardinal directions
J B Haviland · 1998
Earlier work this paper cites.
Towards AI-complete question answering: A set of prerequisite toy tasks
J Weston, A Bordes, S Chopra, A M Rush, B Van Merriënboer, A Joulin, and T Mikolov · 2016
Earlier work this paper cites.
Language models are few-shot learners
T T Brown et al · 2020
Earlier work this paper cites.
SPARTQA: A textual question answering benchmark for spatial reasoning
R Mirzaee, H Rajaby Faghihi, Q Ning, and P Kordjamshidi · 2021
Earlier work this paper cites.
Faithful reasoning using large language models, 2022
A Creswell and M Shanahan · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
StepGame: A new benchmark for robust multi-hop spatial reasoning in texts
Z Shi, Q Zhang, and A Lipani · 2022
Cited alongside, same era.
An evaluation of ChatGPT-4’s Qualitative Spatial Reasoning Capabilities in RCC-8
A G Cohn · 2023
Cited alongside, same era.
Towards a mechanistic interpretation of multi-step reasoning capabilities of language models
Y Hou, J Li, Y Fei, A Stolfo, W Zhou, G Zeng, A Bosselut, and M Sachan · 2023
Cited alongside, same era.
Towards reasoning in large language models: A survey, 2023
J Huang and K C-C Chang · 2023
Cited alongside, same era.
Geolm: Empowering language models for geospatially grounded language understanding
Zekun Li, Wenxuan Zhou, Yao-Yi Chiang, and Muhao Chen · 2023
Advancing spatial reasoning in large language models: An in-depth evaluation and enhancement using the StepGame benchmark
F Li, D C Hogg, and A G Cohn · 2024
Closest in time.
Reframing spatial reasoning evaluation in language models: A real-world simulation benchmark for qualitative reasoning
F Li, D C Hogg, and A G Cohn · 2024
Closest in time.
OpenAI and Josh Achiam et al · 2024
Closest in time.
STEER: Assessing the economic rationality of large language models
Narun Krishnamurthi Raman, Taylor Lundy, Samuel Joseph Amouyal, Yoav Levine, Kevin Leyton-Brown, and Moshe Tennenholtz · 2024
Closest in time.
Visualization-of-thought elicits spatial reasoning in large language models
W Wu, S Mao, Y Zhang, Y Xia, L Dong, L Cui, and F Wei · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey on prompting techniques in LLMs, 2024
Prabin Bhandari · 2024
Cited alongside, same era.
Rationality Report Cards
K Leyton-Brown
Cited in the paper.
Closest in time.
Evaluating spatial understanding of large language models, 2024
Y Yamada, Y Bao, A K Lampinen, J Kasai, and I Yildirim · 2024
Closest in time.