Fetching the paper…
Reading the bibliography…
Visual perspective-taking (VPT), the ability to understand the viewpoint of another person, enables individuals to anticipate the actions of other people.
The Child’s Conception of Space
Jean Piaget and Bärbel Inhelder · 1956
Earlier work this paper cites.
The development of knowledge about visual perception
John H. Flavell · 1977
Earlier work this paper cites.
Does the autistic child have a “theory of mind”?
Simon Baron-Cohen, Alan M. Leslie, and Uta Frith · 1985
Earlier work this paper cites.
Core knowledge
E. S. Spelke · 2000
Earlier work this paper cites.
Spatial updating in humans
J. M. Loomis · 2003
Earlier work this paper cites.
The development of spatial cognition and reasoning
D. R. Montello · 2005
Earlier work this paper cites.
The Learning Brain: Lessons for Education
Uta Frith and Sarah-Jayne Blakemore · 2006
Earlier work this paper cites.
Mindreading: The cognitive basis of "theory of mind."
I. A. Apperly · 2010
Earlier work this paper cites.
Seeing it their way: Evidence for rapid and involuntary computation of what other people see
Dana Samson, Ian Apperly, Jason Braithwaite, Benjamin Andrews, and Sarah Scott · 2010
Earlier work this paper cites.
The role of perspective taking in children’s understanding of knowledge access
Henrike Moll and Michael Tomasello · 2013
Earlier work this paper cites.
A review of visual perspective taking in autism spectrum disorder
Amy Pearson, Danielle Ropar, and Antonia F de C. Hamilton · 2013
Earlier work this paper cites.
The two forms of visual perspective taking are differently embodied and subserve different spatial prepositions
Klaus Kessler and Konstantina E. Rutherford · 2014
Earlier work this paper cites.
Perspective-taking is spontaneous but not automatic
Cathal O’Grady, Thomas Scott-Phillips, Susannah Lavelle, and Kenny Smith · 2020
Earlier work this paper cites.
Aishwarya Agrawal, Ivana Kajić, Emanuele Bugliarello, Elnaz Davoodi, Anita Gergely, Phil Blunsom, and Aida Nematzadeh · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Visual classification via description from large language models
Sachit Menon and Carl Vondrick · 2022
Earlier work this paper cites.
Visual perspective taking is not automatic in a simplified dot task: Evidence from newly sighted children, primary school children and adults
Paula Rubio-Fernandez, Madeleine Long, Vishakha Shukla, Vrinda Bhatia, and Pawan Sinha · 2022
Earlier work this paper cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Cited alongside, same era.
Scaling laws for generative mixed-modal language models
Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Susan Zhang, Stephen Roller, Naman Goyal, Omer Levy, and Luke Zettlemoyer · 2023
Cited alongside, same era.
Homepage
Moondream AI · 2023
Cited alongside, same era.
Claude 3 family
Anthropic · 2023
Cited alongside, same era.
Investigating prompting techniques for zero-and few-shot visual question answering
Rabiul Awal, Le Zhang, and Aishwarya Agrawal · 2023
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Detecting and preventing hallucinations in large vision language models
Anisha Gunjal, Jihan Yin, and Erhan Bas · 2024
Closest in time.
Vision-language models for medical report generation and visual question answering: A review
Iryna Hartsock and Ghulam Rasool · 2024
Closest in time.
Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna · 2024
Closest in time.
D \ \backslash ’ej \ \backslash a vu memorization in vision-language models
Bargav Jayaraman, Chuan Guo, and Kamalika Chaudhuri · 2024
Closest in time.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinze Bai et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Cited alongside, same era.
Gpt-4v system card
OpenAI · 2023
Cited alongside, same era.
Large language models sensitivity to the order of options in multiple-choice questions, 2023
Pouya Pezeshkpour and Estevam Hruschka · 2023
Cited alongside, same era.
To the cutoff… and beyond? a longitudinal perspective on llm data contamination
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Cited alongside, same era.
Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations
Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji, Zhiyuan Liu, and Maosong Sun · 2023
Cited alongside, same era.
Closest in time.
The 3d-pc: a benchmark for visual perspective taking in humans and machines
Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E Lewis, Zygmunt Pizlo, and Thomas Serre · 2024
Closest in time.
Robotic control via embodied chain-of-thought reasoning
Zawalski Michał, Chen William, Pertsch Karl, Mees Oier, Finn Chelsea, and Levine Sergey · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI et al · 2024
Closest in time.
“picture this from there”: spatial perspective-taking in developmental visuospatial disorder and developmental coordination disorder
Camilla Orefice, Ramona Cardillo, Isabella Lonciari, Leonardo Zoccante, and Irene C Mammarella · 2024
Closest in time.
The neglected tails in vision-language models
Shubham Parashar, Zhiqiu Lin, Tian Liu, Xiangjue Dong, Yanan Li, Deva Ramanan, James Caverlee, and Shu Kong · 2024
Closest in time.
Convoi: Context-aware navigation using vision language models in outdoor and indoor environments, 2024
Adarsh Jagan Sathyamoorthy et al · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team et al · 2024
Closest in time.
Vishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma, Philip HS Torr, Adel Bibi, Samuel Albanie, and Matthias Bethge · 2024
Closest in time.
Cogvlm: Visual expert for pretrained language models, 2024
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Closest in time.
Fool your (vision and) language model with embarrassingly simple permutations, 2024
Yongshuo Zong, Tingyang Yu, Ruchika Chavhan, Bingchen Zhao, and Timothy Hospedales · 2024
Closest in time.