Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) exhibit impressive performance across various visual tasks.
Reinforcement learning: An introduction , volume 1
Sutton, R. S.; Barto, A. G.; et al. 1998 · 1998
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
pix2code: Generating code from a graphical user interface screenshot
Beltramelli, T. 2018 · 2018
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L.; Lu, K.; Rajeswaran, A.; Lee, K.; Grover, A.; Laskin, M.; Abbeel, P.; Srinivas, A.; and Mordatch, I. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Earlier work this paper cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W.; Li, Z.; Zhuang, S.; Sheng, Y.; Zheng, L.; Yu, C. H.; Gonzalez, J. E.; Zhang, H.; and Stoica, I. 2023 · 2023
Earlier work this paper cites.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. 2023 · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools, 2023
Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023 · 2023
Earlier work this paper cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C.; Yin, S.; Qi, W.; Wang, X.; Tang, Z.; and Duan, N. 2023 · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023 · 2023
Earlier work this paper cites.
Multimodal chain-of-thought reasoning in language models
Zhang, Z.; Zhang, A.; Li, M.; Zhao, H.; Karypis, G.; and Smola, A. 2023 · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023 · 2023
Cited alongside, same era.
Unlocking the conversion of web screenshots into html code with the websight dataset
Laurençon, H.; Tronchon, L.; and Sanh, V. 2024 · 2024
Cited alongside, same era.
Llava-onevision: Easy visual task transfer
Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; et al. 2024 · 2024
Cited alongside, same era.
Layoutllm: Layout instruction tuning with large language models for document understanding
Luo, C.; Shen, Y.; Zhu, Z.; Zheng, Q.; Yu, Z.; and Yao, C. 2024 · 2024
Cited alongside, same era.
Cogcom: Train large vision-language models diving into details through chain of manipulations
Qi, J.; Ding, M.; Wang, W.; Bai, Y.; Lv, Q.; Hong, W.; Xu, B.; Hou, L.; Li, J.; Dong, Y.; et al. 2024 · 2024
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Huang, W.; Jia, B.; Zhai, Z.; Cao, S.; Ye, Z.; Zhao, F.; Xu, Z.; Hu, Y.; and Lin, S. 2025 · 2025
Closest in time.
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
Mao, S.; Chen, Y.; Cai, P.; Wang, D.; Yan, G.; Yu, Z.; and Shi, B. 2025 · 2025
Closest in time.
gpt-4o-and-more-tools-to-chatgpt-free
OpenAI. 2024a · 2025
Closest in time.
introducing-openai-o1-preview
OpenAI. 2024b · 2025
Closest in time.
Introducing-o3-and-o4-mini
OpenAI. 2025a · 2025
Closest in time.
Thinking with images
OpenAI. 2025b · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chartmimic: Evaluating lmm’s cross-modal reasoning capability via chart-to-code generation
Yang, C.; Shi, C.; Liu, Y.; Shui, B.; Wang, J.; Jing, M.; Xu, L.; Zhu, X.; Li, S.; Zhang, Y.; et al. 2024 · 2024
Cited alongside, same era.
Claude 4
Anthropic. 2025 · 2025
Cited alongside, same era.
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; Zhong, H.; Zhu, Y.; Yang, M.; Li, Z.; Wan, J.; Wang, P.; Ding, W.; Fu, Z.; Xu, Y.; Ye, J.; Zhang, X.; Xie, T.; Cheng, Z.; Zhang, H.; Yang, Z.; Xu, H.; and Lin, J. 2025 · 2025
Cited alongside, same era.
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Chu, T.; Zhai, Y.; Yang, J.; Tong, S.; Xie, S.; Schuurmans, D.; Le, Q. V.; Levine, S.; and Ma, Y. 2025 · 2025
Cited alongside, same era.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI; Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; and et al. 2025 · 2025
Cited alongside, same era.
Refocus: Visual editing as a chain of thought for structured image understanding
Fu, X.; Liu, M.; Yang, Z.; Corring, J.; Lu, Y.; Yang, J.; Roth, D.; Florencio, D.; and Zhang, C. 2025 · 2025
Cited alongside, same era.
Gemini-2-5-model-family
Google. 2025 · 2025
Cited alongside, same era.
The Asymmetry of Verification, and Verifier’s Law
Wei, J. 2025 · 2025
Closest in time.
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
Wu, M.; Yang, J.; Jiang, J.; Li, M.; Yan, K.; Yu, H.; Zhang, M.; Zhai, C.; and Nahrstedt, K. 2025 · 2025
Closest in time.
Xiao, T.; Xu, X.; Huang, Z.; Gao, H.; Liu, Q.; Liu, Q.; and Chen, E. 2025 · 2025
Closest in time.
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
Xu, C.; Wang, Y.; Wei, L.; Sun, L.; and Huang, W. 2025 · 2025
Closest in time.
Chartcoder: Advancing multimodal large language model for chart-to-code generation
Zhao, X.; Luo, X.; Shi, Q.; Chen, C.; Wang, S.; Liu, Z.; and Sun, M. 2025 · 2025
Closest in time.
DeepEyes: Incentivizing “Thinking with Images” via Reinforcement Learning
Zheng, Z.; Yang, M.; Hong, J.; Zhao, C.; Xu, G.; Yang, L.; Shen, C.; and Yu, X. 2025 · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models
Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; Su, W.; Shao, J.; et al. 2025 · 2025
Closest in time.