Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain.
Recovering intrinsic scene characteristics from images
Harry G. Barrow and Jay M. Tenenbaum · 1978
Earlier work this paper cites.
Perceptions as hypotheses
Richard Langton Gregory · 1980
Earlier work this paper cites.
Vision as bayesian inference: analysis by synthesis?
Alan L. Yuille and Daniel Kersten · 2006
Earlier work this paper cites.
Analysis by synthesis: A (re-)emerging program of research for language and vision
Thomas G. Bever and David Poeppel · 2010
Earlier work this paper cites.
Efficient and robust analysis-by-synthesis in vision: A computational framework, behavioral tests, and modeling neuronal representations
Ilker Yildirim, Tejas D Kulkarni, Winrich A Freiwald, and Joshua B Tenenbaum · 2015
Earlier work this paper cites.
Picture: A probabilistic programming language for scene perception
Tejas D. Kulkarni, Pushmeet Kohli, Joshua B. Tenenbaum, and Vikash K. Mansinghka · 2015
Earlier work this paper cites.
Blender - a 3D modelling and rendering package, 2016
Blender Online Community · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap · 2017
Earlier work this paper cites.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Earlier work this paper cites.
Learning to infer graphics programs from hand-drawn images
Kevin Ellis, Daniel Ritchie, Armando Solar-Lezama, and Josh Tenenbaum · 2018
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning · 2018
Earlier work this paper cites.
Multi-target embodied question answering
Licheng Yu, Xinlei Chen, Georgia Gkioxari, Mohit Bansal, Tamara L Berg, and Dhruv Batra · 2019
Earlier work this paper cites.
Soft rasterizer: A differentiable renderer for image-based 3d reasoning
Shichen Liu, Tianye Li, Weikai Chen, and Hao Li · 2019
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
Dave Zhenyu Chen, Angel X. Chang, and Matthias Nießner · 2019
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2019
Earlier work this paper cites.
Learning to infer and execute 3d shape programs
Yonglong Tian, Andrew Luo, Xingyuan Sun, Kevin Ellis, William T Freeman, Joshua B Tenenbaum, and Jiajun Wu · 2019
Earlier work this paper cites.
Program-guided image manipulators
Jiayuan Mao, Xiuming Zhang, Yikai Li, William T Freeman, Joshua B Tenenbaum, and Jiajun Wu · 2019
Earlier work this paper cites.
Monet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Earlier work this paper cites.
Genesis: Generative scene inference and sampling with object-centric latent representations
Martin Engelcke, Adam R Kosiorek, Oiwi Parker Jones, and Ingmar Posner · 2019
Earlier work this paper cites.
Difftaichi: Differentiable programming for physical simulation
Yuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun, Nathan A. Carr, Jonathan Ragan-Kelley, and Frédo Durand · 2020
Earlier work this paper cites.
Differentiable vector graphics rasterization for editing and learning
Tzu-Mao Li, Michal Lukáč, Michaël Gharbi, and Jonathan Ragan-Kelley · 2020
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Earlier work this paper cites.
Synsin: End-to-end view synthesis from a single image
Olivia Wiles, Georgia Gkioxari, and Noah Snavely · 2020
Earlier work this paper cites.
Path-space differentiable rendering
Cheng Zhang, Yihang Guo, Zexiang Dong, Ravi Ramamoorthi, and Manmohan Chandraker · 2020
Earlier work this paper cites.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas J. Guibas · 2020
Earlier work this paper cites.
Scan2cap: Context-aware dense captioning in rgb-d scans
Dave Zhenyu Chen, Ali Gholami, Matthias Nießner, and Angel X. Chang · 2020
Earlier work this paper cites.
Shapeassembly: Learning to generate programs for 3d shape structure synthesis
R Kenny Jones, Theresa Barton, Xianghao Xu, Kai Wang, Ellen Jiang, Paul Guerrero, Niloy J Mitra, and Daniel Ritchie · 2020
Earlier work this paper cites.
Embodied bert: A transformer model for embodied, language-guided visual task completion
Alessandro Suglia, Qiaozi Gao, Jesse Thomason, Govind Thattai, and Gaurav Sukhatme · 2021
Cited alongside, same era.
Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction
Michael Oechsle, Songyou Peng, and Andreas Geiger · 2021
Cited alongside, same era.
Volume rendering of neural implicit surfaces
Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman · 2021
Cited alongside, same era.
Scanqa: 3d question answering for spatial scene understanding
Daich Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoaki Kawanabe · 2021
Cited alongside, same era.
Giraffe: Representing scenes as compositional generative neural feature fields
Michael Niemeyer and Andreas Geiger · 2021
Cited alongside, same era.
https://github.com/modelcontextprotocol , 2024
Model Context Protocol · 2024
Later among the works it cites.
Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness
Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang, and Xihui Liu · 2024
Later among the works it cites.
Openeqa: Embodied question answering in the era of foundation models
Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma, Vincent-Pierre Berges, Shiqi Zhang, Pulkit Agrawal, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton, Alexander Sax, and Aravind Rajeswaran · 2024
Later among the works it cites.
Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang, Xuesong Niu, Tengyu Liu, Qing Li, and Siyuan Huang · 2024
Later among the works it cites.
Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning
Qiao Gu, Ali Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallabhula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty Ellis, Rama Chellappa, Chuang Gan, Celso Miguel de Melo, Joshua B. Tenenbaum, Antonio Torralba, Florian Shkurti, and Liam Paull · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi · 2022
Cited alongside, same era.
One step at a time: Long-horizon vision-and-language navigation with milestones
Chan Hee Song, Jihyung Kil, Tai-Yu Pan, Brian M Sadler, Wei-Lun Chao, and Yu Su · 2022
Cited alongside, same era.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, F. Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R. Florence, and Andy Zeng · 2022
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Cited alongside, same era.
Vipergpt: Visual inference via python execution for reasoning
Didac Suris, Sachit Menon, and Carl Vondrick · 2023
Cited alongside, same era.
Later among the works it cites.
Agent3d-zero: An agent for zero-shot 3d understanding
Sha Zhang, Di Huang, Jiajun Deng, Shixiang Tang, Wanli Ouyang, Tong He, and Yanyong Zhang · 2024
Later among the works it cites.
3dmit: 3d multi-modal instruction tuning for scene understanding
Zeju Li, Chao Zhang, Xiaoyan Wang, et al · 2024
Later among the works it cites.
Openeqa: Embodied question answering in the era of foundation models
Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, et al · 2024
Later among the works it cites.
Scenecraft: An llm agent for synthesizing 3d scenes as blender code
Ziniu Hu, Ahmet Iscen, Aashi Jain, Thomas Kipf, Yisong Yue, David A Ross, Cordelia Schmid, and Alireza Fathi · 2024
Later among the works it cites.
The scene language: Representing scenes with programs, words, and embeddings
Yunzhi Zhang, Zizhang Li, Matt Zhou, Shangzhe Wu, and Jiajun Wu · 2024
Later among the works it cites.
Scenemotifcoder: Example-driven visual program learning for generating 3d object arrangements
Hou In Ivan Tam, Hou In Derek Pun, Austin T Wang, Angel X Chang, and Manolis Savva · 2024
Later among the works it cites.
L3go: Language agents with chain-of-3d-thoughts for generating unconventional objects
Yutaro Yamada, Khyathi Chandu, Yuchen Lin, Jack Hessel, Ilker Yildirim, and Yejin Choi · 2024
Later among the works it cites.
Scenex: Procedural controllable large-scale scene generation via large-language models
Mengqi Zhou, Jun Hou, Chuanchen Luo, Yuxi Wang, Zhaoxiang Zhang, and Junran Peng · 2024
Later among the works it cites.
Layoutvlm: Differentiable optimization of 3d layout via vision-language models
Fan-Yun Sun, Weiyu Liu, Siyi Gu, Dylan Lim, Goutam Bhat, Federico Tombari, Manling Li, Nick Haber, and Jiajun Wu · 2024
Later among the works it cites.
Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al · 2024
Later among the works it cites.
Llava-next: A strong zero-shot video understanding model, 2024
Yuanhan Zhang, Bo Li, haotian Liu, Yong Jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li · 2024
Later among the works it cites.
Meta AI · 2024
Later among the works it cites.
H2ovl-mississippi vision language models technical report
Shaikat M. Galib, Shanshan Wang, Guanshuo Xu, Pascal Pfeiffer, Ryan Chesler, Mark Landry, and SriSatish Ambati · 2024
Later among the works it cites.
Pravesh Agrawal et al · 2024
Later among the works it cites.
Aria: An open multimodal native mixture-of-experts model
Dongxu Li, Yudong Liu, Haoning Wu, Yue Wang, Zhiqi Shen, Bowen Qu, Xinyao Niu, Guoyin Wang, Bei Chen, and Junnan Li · 2024
Later among the works it cites.
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al · 2024
Later among the works it cites.
Qwen An Yang, Baosong Yang, and Beichen Zhang et al · 2024
Later among the works it cites.
gpt4o, 2024
OpenAI · 2024
Later among the works it cites.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al · 2025
Closest in time.
Blendermcp
Siddharth Ahuja · 2025
Closest in time.
Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning, 2025
Hanxun Yu, Wentong Li, Song Wang, Junbo Chen, and Jianke Zhu · 2025
Closest in time.
Spatialrgpt: Grounded spatial reasoning in vision-language models
An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu · 2025
Closest in time.
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Phi-4 Research Team · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Yuchen Duan, Hao Tian, Weijie Su, Jie Shao, et al · 2025
Closest in time.