Fetching the paper…
Reading the bibliography…
Mathematical error detection in educational settings presents a significant challenge for Multimodal Large Language Models (MLLMs), requiring a sophisticated understanding of both visual and textual mathematical content along with complex reasoning capabilities.
Teaching and learning mathematics through error analysis
Sheryl J Rushton. 2018 · 2018
Earlier work this paper cites.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu. 2021 · 2021
Earlier work this paper cites.
Improving image captioning descriptiveness by ranking and llm-based fusion
Simone Bianco, Luigi Celona, Marco Donzella, and Paolo Napoletano. 2023 · 2023
Earlier work this paper cites.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2023 · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2023 · 2023
Earlier work this paper cites.
Multimodal large language models: A survey
Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, and S Yu Philip. 2023a · 2023
Earlier work this paper cites.
Graph of logic: Enhancing llm reasoning with graphs and symbolic logic
Fatimah Alotaibi, Adithya Kulkarni, and Dawei Zhou. 2024 · 2024
Earlier work this paper cites.
Make your llm fully utilize the context
Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, Jian-Guang Lou, and Weizhu Chen. 2024 · 2024
Earlier work this paper cites.
Claude 3.5
Anthropic. 2024 · 2024
Earlier work this paper cites.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024 · 2024
Earlier work this paper cites.
Bofei Gao, Zefan Cai, Runxin Xu, Peiyi Wang, Ce Zheng, Runji Lin, Keming Lu, Dayiheng Liu, Chang Zhou, Wen Xiao, et al. 2024 · 2024
Earlier work this paper cites.
Vtimellm: Empower llm to grasp video moments
Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. 2024 · 2024
Earlier work this paper cites.
Mmneuron: Discovering neuron-level domain-specific interpretation in multimodal large language model
Jiahao Huo, Yibo Yan, Boren Hu, Yutao Yue, and Xuming Hu. 2024 · 2024
Earlier work this paper cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024 · 2024
Earlier work this paper cites.
Llms can find mathematical reasoning mistakes by pedagogical chain-of-thought
Zhuoxuan Jiang, Haoyuan Peng, Shanshan Feng, Fan Li, and Dongsheng Li. 2024 · 2024
Earlier work this paper cites.
Multimodal alignment and fusion: A survey
Songtao Li and Hao Tang. 2024 · 2024
Earlier work this paper cites.
Improving automatic vqa evaluation using large language models
Oscar Mañas, Benno Krojer, and Aishwarya Agrawal. 2024 · 2024
Earlier work this paper cites.
The landscape of emerging ai agent architectures for reasoning, planning, and tool calling: A survey
Tula Masterman, Sandi Besen, Mason Sawtell, and Alex Chao. 2024 · 2024
Earlier work this paper cites.
Aios: Llm agent operating system
Kai Mei, Xi Zhu, Wujiang Xu, Wenyue Hua, Mingyu Jin, Zelong Li, Shuyuan Xu, Ruosong Ye, Yingqiang Ge, and Yongfeng Zhang. 2024 · 2024
Earlier work this paper cites.
Orca-math: Unlocking the potential of slms in grade school math
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah. 2024 · 2024
Cited alongside, same era.
Minheng Ni, Yutao Fan, Lei Zhang, and Wangmeng Zuo. 2024 · 2024
Cited alongside, same era.
GPT-4o system card
OpenAI. 2024 · 2024
Cited alongside, same era.
A review on large language models: Architectures, applications, taxonomies, open issues and challenges
Mohaimenul Azam Khan Raiaan, Md Saddam Hossain Mukta, Kaniz Fatema, Nur Mohammad Fahad, Sadman Sakib, Most Marufatul Jannat Mim, Jubaer Ahmad, Mohammed Eunus Ali, and Sami Azam. 2024 · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Llm agents for education: Advances and applications
Zhendong Chu, Shen Wang, Jian Xie, Tinghui Zhu, Yibo Yan, Jinheng Ye, Aoxiao Zhong, Xuming Hu, Jing Liang, Philip S Yu, et al. 2025 · 2025
Closest in time.
The danger of overthinking: Examining the reasoning-action dilemma in agentic tasks
Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, et al. 2025 · 2025
Closest in time.
Following the autoregressive nature of llm embeddings via compression and alignment
Jingcheng Deng, Zhongtao Jiang, Liang Pang, Liwei Chen, Kun Xu, Zihao Wei, Huawei Shen, and Xueqi Cheng. 2025 · 2025
Closest in time.
Virgo: A preliminary exploration on reproducing o1-like mllm
Yifan Du, Zikang Liu, Yifan Li, Wayne Xin Zhao, Yuqi Huo, Bingning Wang, Weipeng Chen, Zheng Liu, Zhongyuan Wang, and Ji-Rong Wen. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024 · 2024
Cited alongside, same era.
Survey of different large language model architectures: Trends, benchmarks, and challenges
Minghao Shao, Abdul Basit, Ramesh Karri, and Muhammad Shafique. 2024 · 2024
Cited alongside, same era.
Can large language models act as symbolic reasoners?
Rob Sullivan and Nelly Elsayed. 2024 · 2024
Cited alongside, same era.
Introducing qwen1.5
Qwen Team. 2024 · 2024
Cited alongside, same era.
Semantic alignment for multimodal large language models
Tao Wu, Mengze Li, Jingyuan Chen, Wei Ji, Wang Lin, Jinyang Gao, Kun Kuang, Zhou Zhao, and Fei Wu. 2024 · 2024
Cited alongside, same era.
Renqiu Xia, Song Mao, Xiangchao Yan, Hongbin Zhou, Bo Zhang, Haoyang Peng, Jiahao Pi, Daocheng Fu, Wenjie Wu, Hancheng Ye, et al. 2024 · 2024
Cited alongside, same era.
Large multimodal agents: A survey
Junlin Xie, Zhihong Chen, Ruifei Zhang, Xiang Wan, and Guanbin Li. 2024 · 2024
Cited alongside, same era.
Building math agents with multi-turn iterative preference learning
Wei Xiong, Chengshuai Shi, Jiaming Shen, Aviv Rosenberg, Zhen Qin, Daniele Calandriello, Misha Khalman, Rishabh Joshi, Bilal Piot, Mohammad Saleh, et al. 2024 · 2024
Cited alongside, same era.
Jiahao Huo, Yibo Yan, Xu Zheng, Yuanhuiyi Lyu, Xin Zou, Zhihua Wei, and Xuming Hu. 2025 · 2025
Closest in time.
On opportunities and challenges of large multimodal foundation models in education
Stefan Küchemann, Karina E Avila, Yavuz Dinc, Chiara Hortmann, Natalia Revenga, Verena Ruf, Niklas Stausberg, Steffen Steinert, Frank Fischer, Martin Fischer, et al. 2025 · 2025
Closest in time.
Step-by-step correction of llm-based math word problems solutions
Yiyao Li, Dhanish Musharraf Ubaidali, Lu Wang, and Wenyu Zhang. 2025a · 2025
Closest in time.
From alt-text to real context: Revolutionizing image captioning using the potential of llm
Vrajkumar Patel, Aayush Modi, Harsh Mistry, Abhishesh Mishra, Rocky Upadhyay, and Apoorva Shah. 2025 · 2025
Closest in time.
“mathematics education in the era of chatgpt: Investigating its meaning and use for school and university education”—editorial to special issue
Birgit Pepin, Nils Buchholtz, and Ulises Salinas-Hernández. 2025 · 2025
Closest in time.
Prmbench: A fine-grained and challenging benchmark for process-level reward models
Mingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou, and Yu Cheng. 2025 · 2025
Closest in time.
Hallucinations of large multimodal models: Problem and countermeasures
Shiliang Sun, Zhilin Lin, and Xuhan Wu. 2025 · 2025
Closest in time.
Llamav-o1: Rethinking step-by-step visual reasoning in llms
Omkar Thawakar, Dinura Dissanayake, Ketan More, Ritesh Thawkar, Ahmed Heakl, Noor Ahsan, Yuhao Li, Mohammed Zumri, Jean Lahoud, Rao Muhammad Anwer, et al. 2025 · 2025
Closest in time.
Videoqa in the era of llms: An empirical study
Junbin Xiao, Nanxin Huang, Hangyu Qin, Dongyang Li, Yicong Li, Fengbin Zhu, Zhulin Tao, Jianxing Yu, Liang Lin, Tat-Seng Chua, et al. 2025 · 2025
Closest in time.
Towards large reasoning models: A survey of reinforced reasoning with large language models
Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al. 2025 · 2025
Closest in time.
Position: Multimodal large language models can significantly advance scientific reasoning
Yibo Yan, Shen Wang, Jiahao Huo, Jingheng Ye, Zhendong Chu, Xuming Hu, Philip S Yu, Carla Gomes, Bart Selman, and Qingsong Wen. 2025 · 2025
Closest in time.
Position: Llms can be good tutors in foreign language education
Jingheng Ye, Shen Wang, Deqing Zou, Yibo Yan, Kun Wang, Hai-Tao Zheng, Zenglin Xu, Irwin King, Philip S Yu, and Qingsong Wen. 2025 · 2025
Closest in time.
A survey of multimodal learning: Methods, applications, and future
Yuan Yuan, Zhaojian Li, and Bin Zhao. 2025 · 2025
Closest in time.
From correctness to comprehension: Ai agents for personalized error diagnosis in education
Yi-Fan Zhang, Hang Li, Dingjie Song, Lichao Sun, Tianlong Xu, and Qingsong Wen. 2025 · 2025
Closest in time.
R1-omni: Explainable omni-multimodal emotion recognition with reinforcing learning
Jiaxing Zhao, Xihan Wei, and Liefeng Bo. 2025 · 2025
Closest in time.