Fetching the paper…
Reading the bibliography…
Visual transformation reasoning (VTR) is a vital cognitive capability that empowers intelligent agents to understand dynamic scenes, model causal relationships, and predict future states, and thereby guiding actions and laying the foundation for advanced intelligent systems.
The construction of reality in the child
Piaget, J. 2013 · 2013
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J.; Hariharan, B.; Van Der Maaten, L.; Fei-Fei, L.; Lawrence Zitnick, C.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
Cooking with blocks: A recipe for visual reasoning on image-pairs
Gokhale, T.; Sampat, S.; Fang, Z.; Yang, Y.; and Baral, C. 2019 · 2019
Earlier work this paper cites.
Robust change captioning
Park, D. H.; et al. 2019 · 2019
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; and Zhou, J. 2019 · 2019
Earlier work this paper cites.
Transformation driven visual reasoning
Hong, X.; Lan, Y.; Pang, L.; Guo, J.; and Cheng, X. 2021 · 2021
Earlier work this paper cites.
Visual reasoning: From state to transformation
Hong, X.; Lan, Y.; Pang, L.; Guo, J.; and Cheng, X. 2023 · 2023
Earlier work this paper cites.
Agrawal, P.; Antoniak, S.; Hanna, E. B.; Bout, B.; Chaplot, D.; Chudnovsky, J.; Costa, D.; De Monicault, B.; Garg, S.; Gervet, T.; et al. 2024 · 2024
Earlier work this paper cites.
Spatialrgpt: Grounded spatial reasoning in vision-language models
Cheng, A.-C.; Yin, H.; Fu, Y.; Guo, Q.; Yang, R.; Kautz, J.; Wang, X.; and Liu, S. 2024 · 2024
Earlier work this paper cites.
The development of human causal learning and reasoning
Goddu, M. K.; and Gopnik, A. 2024 · 2024
Earlier work this paper cites.
Latency-aware unified dynamic networks for efficient image recognition
Han, Y.; Liu, Z.; Yuan, Z.; Pu, Y.; Wang, C.; Song, S.; and Huang, G. 2024 · 2024
Earlier work this paper cites.
Prompting large language model with context and pre-answer for knowledge-based VQA
Hu, Z.; Yang, P.; Jiang, Y.; and Bai, Z. 2024 · 2024
Earlier work this paper cites.
What foundation models can bring for robot learning in manipulation: A survey
Li, D.; Jin, Y.; Sun, Y.; Yu, H.; Shi, J.; Hao, X.; Hao, P.; Liu, H.; Sun, F.; Zhang, J.; et al. 2024 · 2024
Earlier work this paper cites.
Vila: On pre-training for visual language models
Lin, J.; Yin, H.; Ping, W.; Molchanov, P.; Shoeybi, M.; and Han, S. 2024 · 2024
Earlier work this paper cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Jiang, Q.; Li, C.; Yang, J.; Su, H.; et al. 2024 · 2024
Earlier work this paper cites.
ChartAssistant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
Meng, F.; Shao, W.; Lu, Q.; Gao, P.; Zhang, K.; Qiao, Y.; and Luo, P. 2024 · 2024
Cited alongside, same era.
Llama3.2-Vision Model Card
Meta. 2025b · 2024
Cited alongside, same era.
SAT: Spatial Aptitude Training for Multimodal Language Models
Ray, A.; Duan, J.; Tan, R.; Bashkirova, D.; Hendrix, R.; Ehsani, K.; Kembhavi, A.; Plummer, B. A.; Krishna, R.; Zeng, K.-H.; and Saenko, K. 2024 · 2024
Cited alongside, same era.
Dettoolchain: A new prompting paradigm to unleash detection ability of mllm
Wu, Y.; Wang, Y.; Tang, S.; Wu, W.; He, T.; Ouyang, W.; Torr, P.; and Wu, J. 2024 · 2024
Cited alongside, same era.
Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras
Abouelenin, A.; Ashfaq, A.; Atkinson, A.; Awadalla, H.; Bach, N.; Bao, J.; Benhaim, A.; Cai, M.; Chaudhary, V.; Chen, C.; et al. 2025 · 2025
Guardreasoner-vl: Safeguarding vlms via reinforced reasoning
Liu, Y.; Zhai, S.; Du, M.; Chen, Y.; Cao, T.; Gao, H.; Wang, C.; Li, X.; Wang, K.; Fang, J.; et al. 2025 · 2025
Closest in time.
EgoPrompt: Prompt Pool Learning for Egocentric Action Recognition
Lyu, H.; Chen, C.; Ji, Y.; and Xu, C. 2025 · 2025
Closest in time.
The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation
Meta. 2025a · 2025
Closest in time.
Introducing GPT-4.1 in the API
OpenAI. 2025a · 2025
Closest in time.
OpenAI o3 and o4-mini System Card
OpenAI. 2025b · 2025
Closest in time.
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Claude 3.7 Sonnet and Claude Code
Anthropic. 2025 · 2025
Cited alongside, same era.
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. 2025 · 2025
Cited alongside, same era.
Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
Chen, S.; Zhang, J.; Zhu, T.; Liu, W.; Gao, S.; Xiong, M.; Li, M.; and He, J. 2025 · 2025
Cited alongside, same era.
Gemini 2.5 Pro Preview: even better coding performance
Google. 2025 · 2025
Cited alongside, same era.
GLM-4.1 V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Hong, W.; Yu, W.; Gu, X.; Wang, G.; Gan, G.; Tang, H.; Cheng, J.; Qi, J.; Ji, J.; Pan, L.; et al. 2025 · 2025
Cited alongside, same era.
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
Hoque, R.; Huang, P.; Yoon, D. J.; Sivapurapu, M.; and Zhang, J. 2025 · 2025
Cited alongside, same era.
Foundation models and intelligent decision-making: Progress, challenges, and perspectives
Huang, J.; Xu, Y.; Wang, Q.; Wang, Q. C.; Liang, X.; Wang, F.; Zhang, Z.; Wei, W.; Zhang, B.; Huang, L.; et al. 2025 · 2025
Cited alongside, same era.
Qi, J.; Ding, M.; Wang, W.; Bai, Y.; Lv, Q.; Hong, W.; Xu, B.; Hou, L.; Li, J.; Dong, Y.; and Tang, J. 2025 · 2025
Closest in time.
From Chart to QA Pairs: A Context-Aware Generation Framework for Chart-Containing Documents
Shen, Q.; et al. 2025 · 2025
Closest in time.
Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics
Song, C. H.; Blukis, V.; Tremblay, J.; Tyree, S.; Su, Y.; and Birchfield, S. 2025 · 2025
Closest in time.
Affordgrasp: In-context affordance reasoning for open-vocabulary task-oriented grasping in clutter
Tang, Y.; Zhang, S.; Hao, X.; Wang, P.; Wu, J.; Wang, Z.; and Zhang, S. 2025 · 2025
Closest in time.
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
Wang, Y.; Wang, S.; Cheng, Q.; Fei, Z.; Ding, L.; Guo, Q.; Tao, D.; and Qiu, X. 2025 · 2025
Closest in time.
ViDDAR: Vision language model-based task-detrimental content detection for augmented reality
Xiu, Y.; et al. 2025 · 2025
Closest in time.
Training-free Generation of Temporally Consistent Rewards from VLMs
Zhao, Y.; Yuan, J.; Xu, Z.; Hao, X.; Zhang, X.; Wu, K.; Che, Z.; Liu, C. H.; and Tang, J. 2025 · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models
Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; Su, W.; Shao, J.; et al. 2025 · 2025
Closest in time.
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
Zong, Y.; Zhang, Q.; An, D.; Li, Z.; Xu, X.; Xu, L.; Tu, Z.; Xing, Y.; and Dabeer, O. 2025 · 2025
Closest in time.