Fetching the paper…
Reading the bibliography…
Recent progress in multimodal reasoning has been significantly advanced by textual Chain-of-Thought (CoT), a paradigm where models conduct reasoning within language.
Why a diagram is (sometimes) worth ten thousand words
Jill H Larkin and Herbert A Simon · 1987
Earlier work this paper cites.
The opencv library
Gary Bradski · 2000
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
TRL: Transformer Reinforcement Learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, Quentin Gallouédec, Thomas Wolf, Philipp Schmid, Louis Debut, and Victor Sanh · 2020
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Earlier work this paper cites.
Promptcap: Prompt-guided task-aware image captioning
Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A Smith, and Jiebo Luo · 2022
Earlier work this paper cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, et al · 2022
Earlier work this paper cites.
Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression
Jiaqi Chen, Tong Li, Jinghui Qin, Pan Lu, Liang Lin, Chongyu Chen, and Xiaodan Liang · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
LangChain, 10 2022
Harrison Chase · 2022
Earlier work this paper cites.
An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Jie Cao and Jing Xiao · 2022
Earlier work this paper cites.
LlamaIndex, 11 2022
Jerry Liu · 2022
Earlier work this paper cites.
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2022
Earlier work this paper cites.
Deep learning for deepfakes creation and detection: A survey
Thanh Thi Nguyen, Quoc Viet Hung Nguyen, Dung Tien Nguyen, Duc Thanh Nguyen, Thien Huynh-The, Saeid Nahavandi, Thanh Tam Nguyen, Quoc-Viet Pham, and Cuong M Nguyen · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
Estelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan, and Vasudev Lal · 2022
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2023
Earlier work this paper cites.
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi · 2023
Earlier work this paper cites.
What does clip know about a red circle? visual prompt engineering for vlms
Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi · 2023
Earlier work this paper cites.
Llava-plus: Learning to use tools for creating multimodal agents
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, Jianfeng Gao, and Chunyuan Li · 2023
Earlier work this paper cites.
Multimodal large language models: A survey
Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, and Philip S Yu · 2023
Earlier work this paper cites.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick · 2023
Earlier work this paper cites.
Generating images with multimodal language models
Jing Yu Koh, Daniel Fried, and Russ R Salakhutdinov · 2023
Earlier work this paper cites.
Minigpt-5: Interleaved vision-and-language generation via generative vokens
Kaizhi Zheng, Xuehai He, and Xin Eric Wang · 2023
Earlier work this paper cites.
A multi-modal neural geometric solver with textual clauses parsed from diagram
Ming-Liang Zhang, Fei Yin, and Cheng-Lin Liu · 2023
Earlier work this paper cites.
Geomverse: A systematic evaluation of large models for geometric reasoning
Mehran Kazemi, Hamidreza Alvari, Ankit Anand, Jialin Wu, Xi Chen, and Radu Soricut · 2023
Earlier work this paper cites.
CrewAI: Framework for orchestrating role-playing, autonomous AI agents, 2023
crewAIInc · 2023
Earlier work this paper cites.
Axolotl: Streamlined AI Model Post-Training, 2023
Axolotl AI Cloud Contributors · 2023
Earlier work this paper cites.
Haystack: The Production-Ready Open Source AI Framework, 2023
deepset · 2023
Earlier work this paper cites.
AutoGPT: An Autonomous GPT-4 Experiment, 3 2023
Toran Bruce Richards and Significant Gravitas · 2023
Earlier work this paper cites.
Unsloth: Finetune LLMs 2x faster with up to 70% less memory, 2023
UnslothAI · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model, March 2023
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence · 2023
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts, May 2023
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan · 2023
Earlier work this paper cites.
Geochat: Grounded large vision-language model for remote sensing, November 2023
Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan · 2023
Earlier work this paper cites.
V*: Guided visual search as a core mechanism in multimodal llms
Penghao Wu and Saining Xie · 2023
Earlier work this paper cites.
Pal: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Earlier work this paper cites.
Diffusion models in vision: A survey
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah · 2023
Earlier work this paper cites.
Analyzing and mitigating object hallucination in large vision-language models
Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Earlier work this paper cites.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen · 2023
Earlier work this paper cites.
Vbench: Comprehensive benchmark suite for video generative models, 2023
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu · 2023
Earlier work this paper cites.
Chartbench: A benchmark for complex visual reasoning in charts
Zhengzhuo Xu, Sinan Du, Yiyan Qi, Chengjin Xu, Chun Yuan, and Jian Guo · 2023
Earlier work this paper cites.
M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models
Wenxuan Zhang, Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing · 2023
Earlier work this paper cites.
FlowiseAI: Build AI Agents Visually, 2023
FlowiseAI · 2023
Earlier work this paper cites.
AutoChain: Build lightweight, extensible, and testable LLM Agents, 2023
Forethought Technologies · 2023
Earlier work this paper cites.
trlX: A framework for large scale reinforcement learning from human feedback
Alexander Havrilla, Maksym Zhuravinskyi, Duy Phung, Aman Tiwari, Jonathan Tow, Stella Biderman, Quentin Anthony, and Louis Castricato · 2023
Earlier work this paper cites.
Med-flamingo: a multimodal medical few-shot learner
Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar · 2023
Earlier work this paper cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansfield, Dina Demner-Fushman, Blaise Agüera y Arcas, Dale Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan · 2023
Earlier work this paper cites.
Sparks of large audio models: A survey and outlook
Siddique Latif, Moazzam Shoukat, Fahad Shamshad, Muhammad Usama, Yi Ren, Heriberto Cuayáhuitl, Wenwu Wang, Xulong Zhang, Roberto Togneri, Erik Cambria, et al · 2023
Earlier work this paper cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2023
Cited alongside, same era.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al · 2023
Cited alongside, same era.
Synthetic vision: Training vision-language models to understand physics, December 2024
Vahid Balazadeh, Mohammadmehdi Ataei, Hyunmin Cheong, Amir Hosein Khasahmadi, and Rahul G. Krishnan · 2024
Cited alongside, same era.
Uncovering bias in large vision-language models at scale with counterfactuals
Phillip Howard, Kathleen C Fraser, Anahita Bhiwandiwalla, and Svetlana Kiritchenko · 2024
Later among the works it cites.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2024
Later among the works it cites.
Satori-r1: Incentivizing multimodal reasoning with spatial grounding and verifiable rewards
Chuming Shen, Wei Wei, Xiaoye Qu, and Yu Cheng · 2025
Closest in time.
Explorer: Scaling exploration-driven web trajectory synthesis for multimodal web agents
Vardaan Pahuja, Yadong Lu, Corby Rosset, Boyu Gou, Arindam Mitra, Spencer Whitehead, Yu Su, and Ahmed Awadallah · 2025
Closest in time.
Visual thoughts: A unified perspective of understanding multimodal chain-of-thought
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, and Jianwei Yin · 2024
Cited alongside, same era.
Taco: Learning multi-modal action models with synthetic chains-of-thought-and-action, 2024
Zixian Ma, Jianguo Zhang, Zhiwei Liu, Jieyu Zhang, Juntao Tan, Manli Shu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Caiming Xiong, Ranjay Krishna, and Silvio Savarese · 2024
Cited alongside, same era.
Cogcom: Train large vision-language models diving into details through chain of manipulations
Ji Qi, Ming Ding, Weihan Wang, Yushi Bai, Qingsong Lv, Wenyi Hong, Bin Xu, Lei Hou, Juanzi Li, Yuxiao Dong, and Jie Tang · 2024
Cited alongside, same era.
Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models, 2024
Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li · 2024
Cited alongside, same era.
Cad-assistant: Tool-augmented vllms as generic cad task solvers?
Dimitrios Mallis, Ahmet Serdar Karadeniz, Sebastian Cavada, Danila Rukhovich, Niki Foteinopoulou, Kseniya Cherenkova, Anis Kacem, and Djamila Aouada · 2024
Cited alongside, same era.
Sketchagent: Language-driven sequential sketch generation
Yael Vinker, Tamar Rott Shaham, Kristine Zheng, Alex Zhao, Judith E Fan, and Antonio Torralba · 2024
Cited alongside, same era.
Mmfactory: A universal solution search engine for vision-language tasks
Wan-Cyuan Fan, Tanzila Rahman, and Leonid Sigal · 2024
Cited alongside, same era.
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team · 2024
Cited alongside, same era.
Zihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang, Weiyun Wang, Hao Fei, Yidong Wang, Alex Jinpeng Wang, Zhi Chen, Wanxiang Che, et al · 2025
Closest in time.
Don’t look only once: Towards multimodal interactive reasoning with selective visual revisitation
Jiwan Chung, Junhyeok Kim, Siyeol Kim, Jaeyoung Lee, Min Soo Kim, and Youngjae Yu · 2025
Closest in time.
One rl to see them all: Visual triple unified reinforcement learning
Yan Ma, Linge Du, Xuyang Shen, Shaoxiang Chen, Pengfei Li, Qibing Ren, Lizhuang Ma, Yuchao Dai, Pengfei Liu, and Junjie Yan · 2025
Closest in time.
Point-rft: Improving multimodal reasoning with visually grounded reinforcement finetuning, 2025
Minheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin, Kevin Lin, Wangmeng Zuo, and Lijuan Wang · 2025
Closest in time.
Deepeyes: Incentivizing "thinking with images" via reinforcement learning, 2025
Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu · 2025
Closest in time.
Scaling text-rich image understanding via code-guided synthetic multimodal data generation
Yue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta, Luca Weihs, Andrew Head, Mark Yatskar, Chris Callison-Burch, Ranjay Krishna, Aniruddha Kembhavi, et al · 2025
Closest in time.
Rongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang, Hao Li, Shilin Yan, Hao Tian, Xingyu Zeng, Rui Zhao, Jifeng Dai, Xihui Liu, and Hongsheng Li · 2025
Closest in time.
Visual planning: Let’s think only with images, 2025
Yi Xu, Chengzu Li, Han Zhou, Xingchen Wan, Caiqi Zhang, Anna Korhonen, and Ivan Vulić · 2025
Closest in time.
Got-r1: Unleashing reasoning capability of mllm for visual generation with reinforcement learning
Chengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang, Linjiang Huang, Xingyu Zeng, Hongsheng Li, and Xihui Liu · 2025
Closest in time.
ROLL Team and Other ROLL Contributors · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al · 2025
Closest in time.
Robotic control via embodied chain-of-thought reasoning, March 2025
Michał Zawalski, William Chen, Karl Pertsch, Oier Mees, Chelsea Finn, and Sergey Levine · 2025
Closest in time.
Chemmllm: Chemical multimodal large language model
Qian Tan, Dongzhan Zhou, Peng Xia, Wanhao Liu, Wanli Ouyang, Lei Bai, Yuqiang Li, and Tianfan Fu · 2025
Closest in time.
Tianwei Lin, Wenqiao Zhang, Sijing Li, Yuqian Yuan, Binhe Yu, Haoyuan Li, Wanggui He, Hao Jiang, Mengze Li, Xiaohui Song, et al · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models, 2025
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and Xia Hu · 2025
Closest in time.
A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond
Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, et al · 2025
Closest in time.
Benchmarking multimodal knowledge conflict for large multimodal models
Yifan Jia, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin, Rui Yang, Fenze Feng, Mingcai Chen, Hengyang Lu, Haozhe Wang, et al · 2025
Closest in time.
Advancing vision-language models in front-end development via data synthesis
Tong Ge, Yashu Liu, Jieping Ye, Tianyi Li, and Chao Wang · 2025
Closest in time.
Transfer between modalities with metaqueries
Xichen Pan, Satya Narayan Shukla, Aashu Singh, Zhuokai Zhao, Shlok Kumar Mishra, Jialiang Wang, Zhiyang Xu, Jiuhai Chen, Kunpeng Li, Felix Juefei-Xu, et al · 2025
Closest in time.
Mogao: An omni foundation model for interleaved multi-modal generation
Chao Liao, Liyang Liu, Xun Wang, Zhengxiong Luo, Xinyu Zhang, Wenliang Zhao, Jie Wu, Liang Li, Zhi Tian, and Weilin Huang · 2025
Closest in time.
Thinking with generated images
Ethan Chern, Zhulin Hu, Steffi Chern, Siqi Kou, Jiadi Su, Yan Ma, Zhijie Deng, and Pengfei Liu · 2025
Closest in time.
Show-o2: Improved native unified multimodal models
Jinheng Xie, Zhenheng Yang, and Mike Zheng Shou · 2025
Closest in time.
Feng Han, Yang Jiao, Shaoxiang Chen, Junhao Xu, Jingjing Chen, and Yu-Gang Jiang · 2025
Closest in time.
MATHGLANCE: multimodal large language models do not know where to look in mathematical diagrams
Yanpeng Sun, Shan Zhang, Wei Tang, Aotian Chen, Piotr Koniusz, Kai Zou, Yuan Xue, and Anton van den Hengel · 2025
Closest in time.
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gemini Team, Google · 2025
Closest in time.
Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models
Chengke Zou, Xingang Guo, Rui Yang, Junyu Zhang, Bin Hu, and Huan Zhang · 2025
Closest in time.
Explain with visual keypoints like a real mentor! a benchmark for multimodal solution explanation
Jaewoo Park, Jungyang Park, Dongju Jang, Jiwan Chung, Byungwoo Yoo, Jaewoo Shin, Seonjoon Park, Taehyeong Kim, and Youngjae Yu · 2025
Closest in time.
Can mllms reason in multimodality? EMMA: an enhanced multimodal reasoning benchmark
Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, and Yu Cheng · 2025
Closest in time.
Mmscibench: Benchmarking language models on chinese multimodal scientific problems
Xinwu Ye, Chengfan Li, Siming Chen, Wei Wei, and Xiangru Tang · 2025
Closest in time.
Vgrp-bench: Visual grid reasoning puzzle benchmark for large vision-language models
Yufan Ren, Konstantinos Tertikas, Shalini Maiti, Junlin Han, Tong Zhang, Sabine Süsstrunk, and Filippos Kokkinos · 2025
Closest in time.
Are large vision language models good game players?
Xinyu Wang, Bohan Zhuang, and Qi Wu · 2025
Closest in time.
Jing Bi, Junjia Guo, Susan Liang, Guangyu Sun, Luchuan Song, Yunlong Tang, Jinxi He, Jiarui Wu, Ali Vosoughi, Chen Chen, et al · 2025
Closest in time.
Plot2code: A comprehensive benchmark for evaluating multi-modal large language models in code generation from scientific plots
Chengyue Wu, Zhixuan Liang, Yixiao Ge, Qiushan Guo, Zeyu Lu, Jiahao Wang, Ying Shan, and Ping Luo · 2025
Closest in time.
Design2code: Benchmarking multimodal code generation for automated front-end engineering
Chenglei Si, Yanzhe Zhang, Ryan Li, Zhengyuan Yang, Ruibo Liu, and Diyi Yang · 2025
Closest in time.
Sketch2code: Evaluating vision-language models for interactive web design prototyping
Ryan Li, Yanzhe Zhang, and Diyi Yang · 2025
Closest in time.
Tablebench: A comprehensive and complex benchmark for table question answering
Xianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang, Jiaheng Liu, Xeron Du, Di Liang, Daixin Shu, Xianfu Cheng, Tianzhen Sun, Tongliang Li, Zhoujun Li, and Guanglin Niu · 2025
Closest in time.
Multichartqa: Benchmarking vision-language models on multi-chart problems
Zifeng Zhu, Mengzhao Jia, Zhihan Zhang, Lang Li, and Meng Jiang · 2025
Closest in time.
Chartmuseum: Testing visual reasoning capabilities of large vision-language models
Liyan Tang, Grace Kim, Xinyu Zhao, Thom Lake, Wenxuan Ding, Fangcong Yin, Prasann Singhal, Manya Wadhwa, Zeyu Leo Liu, Zayne Sprague, et al · 2025
Closest in time.
Run Hugging Face models instantly with Day-0 support from NVIDIA NeMo Framework
Shashank Verma, Alexandros Koumparoulis, Wenwen Gao, and Bernard Nguyen · 2025
Closest in time.
Position: Multimodal large language models can significantly advance scientific reasoning
Yibo Yan, Shen Wang, Jiahao Huo, Jingheng Ye, Zhendong Chu, Xuming Hu, Philip S Yu, Carla Gomes, Bart Selman, and Qingsong Wen · 2025
Closest in time.
Microvqa: A multimodal reasoning benchmark for microscopy-based scientific research
James Burgess, Jeffrey J Nirschl, Laura Bravo-Sánchez, Alejandro Lozano, Sanket Rajan Gupte, Jesus G Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, et al · 2025
Closest in time.
Visual cognition in multimodal large language models
Luca M. Schulze Buschoff, Elif Akata, Matthias Bethge, and Eric Schulz · 2025
Closest in time.
Llama-nemotron: Efficient reasoning models
Akhiad Bercovich, Itay Levy, Izik Golan, Mohammad Dabbah, Ran El-Yaniv, Omri Puny, Ido Galil, Zach Moshe, Tomer Ronen, Najeeb Nabwani, et al · 2025
Closest in time.
Vila-m3: Enhancing vision-language models with medical expert knowledge
Vishwesh Nath, Wenqi Li, Dong Yang, Andriy Myronenko, Mingxin Zheng, Yao Lu, Zhijian Liu, Hongxu Yin, Yee Man Law, Yucheng Tang, et al · 2025
Closest in time.
Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models
Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, and Xiaofeng Yang · 2025
Closest in time.
Prompting medical large vision-language models to diagnose pathologies by visual question answering
Danfeng Guo and Demetri Terzopoulos · 2025
Closest in time.
How well do llms compress their own chain-of-thought? a token complexity approach
Ayeong Lee, Ethan Che, and Tianyi Peng · 2025
Closest in time.
Prmbench: A fine-grained and challenging benchmark for process-level reward models
Mingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou, and Yu Cheng · 2025
Closest in time.
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen · 2053
Closest in time.