Fetching the paper…
Reading the bibliography…
We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Tall: Temporal activity localization via language query
J. Gao, C. Sun, Z. Yang, and R. Nevatia · 2017
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Generalized intersection over union: A metric and a loss for bounding box regression
H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese · 2019
Earlier work this paper cites.
Websrc: a dataset for web-based structural reading comprehension
X. Chen, Z. Zhao, L. Chen, D. Zhang, J. Ji, A. Luo, Y. Xiong, and K. Yu · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Earlier work this paper cites.
Infographicvqa
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Earlier work this paper cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Earlier work this paper cites.
Egoschema: A diagnostic benchmark for very long-form video language understanding
K. Mangalam, R. Akshulakov, and J. Malik · 2023
Earlier work this paper cites.
Teaching clip to count to ten
R. Paiss, A. Ephrat, O. Tov, S. Zada, I. Mosseri, M. Irani, and T. Dekel · 2023
Earlier work this paper cites.
V*: Guided visual search as a core mechanism in multimodal llms, 2023
P. Wu and S. Xie · 2023
Earlier work this paper cites.
H. Xu, S. Xie, X. E. Tan, P.-Y. Huang, R. Howes, V. Sharma, S.-W. Li, G. Ghosh, L. Zettlemoyer, and C. Feichtenhofer · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Cited alongside, same era.
Instruction-following evaluation for large language models, 2023
J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al · 2023
Cited alongside, same era.
π \pi 0: A vision-language-action flow model for general robot control
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
P. Tong, E. Brown, P. Wu, S. Woo, A. J. V. IYER, S. C. Akula, S. Yang, J. Yang, M. Middepogu, Z. Wang, et al · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen · 2024
Later among the works it cites.
Os-atlas: A foundation action model for generalist gui agents
Z. Wu, Z. Wu, F. Xu, Y. Wang, Q. Sun, C. Jia, K. Cheng, Z. Ding, L. Chen, P. P. Liang, et al · 2024
Later among the works it cites.
Logicvista: Multimodal llm logical reasoning benchmark in visual contexts
Y. Xiao, E. Sun, T. Liu, and W. Wang · 2024
Later among the works it cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky · 2024
Cited alongside, same era.
Seeclick: Harnessing gui grounding for advanced visual gui agents
K. Cheng, Q. Sun, Y. Chu, F. Xu, Y. Li, J. Zhang, and Z. Wu · 2024
Cited alongside, same era.
Visionarena: 230k real world user-vlm conversations with preference labels
C. Chou, L. Dunlap, K. Mashita, K. Mandal, T. Darrell, I. Stoica, J. Gonzalez, and W.-L. Chiang · 2024
Cited alongside, same era.
Nvlm: Open frontier-class multimodal llms
W. Dai, N. Lee, B. Wang, Z. Yang, Z. Liu, J. Barker, T. Rintamaki, M. Shoeybi, B. Catanzaro, and W. Ping · 2024
Cited alongside, same era.
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, et al · 2024
Cited alongside, same era.
C. He, R. Luo, Y. Bai, S. Hu, Z. L. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, et al · 2024
Cited alongside, same era.
Mantis: Interleaved multi-image instruction tuning
D. Jiang, X. He, H. Zeng, C. Wei, M. Ku, Q. Liu, and W. Chen · 2024
Cited alongside, same era.
Prismatic vlms: Investigating the design space of visually-conditioned language models
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Cited alongside, same era.
T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, et al · 2024
Later among the works it cites.
Aguvis: Unified pure vision agents for autonomous gui interaction
Y. Xu, Z. Wang, J. Wang, D. Lu, T. Xie, A. Saha, D. Sahoo, T. Yu, and C. Xiong · 2024
Later among the works it cites.
C. Zou, X. Guo, R. Yang, J. Zhang, B. Hu, and H. Zhang · 2024
Later among the works it cites.
R1-v: Reinforcing super generalization ability in vision-language models with less than $3
L. Chen, L. Li, H. Zhao, Y. Song, and Vinci · 2025
Closest in time.
Supergpqa: Scaling llm evaluation across 285 graduate disciplines
X. Du, Y. Yao, K. Ma, B. Wang, T. Zheng, K. Zhu, M. Liu, Y. Liang, X. Jin, Z. Wei, et al · 2025
Closest in time.
Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos
K. Hu, P. Wu, F. Pu, W. Xiao, Y. Zhang, X. Yue, B. Li, and Z. Liu · 2025
Closest in time.
Tulu 3: Pushing frontiers in open language model post-training, 2025
N. Lambert, J. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, S. Lyu, Y. Gu, S. Malik, V. Graf, J. D. Hwang, J. Yang, R. L. Bras, O. Tafjord, C. Wilhelm, L. Soldaini, N. A. Smith, Y. Wang, P. Dasigi, and H. Hajishirzi · 2025
Closest in time.
Screenspot-pro: Gui grounding for professional high-resolution computer use
K. Li, Z. Meng, H. Lin, Z. Luo, Y. Tian, J. Ma, Z. Huang, and T.-S. Chua · 2025
Closest in time.
American invitational mathematics examination - aime
MAA · 2025
Closest in time.
Computer-using agent: Introducing a universal interface for ai to interact with the digital world
OpenAI · 2025
Closest in time.
Vision language models are blind: Failing to translate detailed visual features into words, 2025
P. Rahmanzadehgervi, L. Bolton, M. R. Taesiri, and A. T. Nguyen · 2025
Closest in time.
Timezero: Temporal video grounding with reasoning-guided lvlm
Y. Wang, B. Xu, Z. Yue, Z. Xiao, Z. Wang, L. Zhang, D. Yang, W. Wang, and Q. Jin · 2025
Closest in time.
Mimo: Unlocking the reasoning potential of language model–from pretraining to posttraining
L.-C.-T. Xiaomi · 2025
Closest in time.
Scaling computer-use grounding via user interface decomposition and synthesis, 2025
T. Xie, J. Deng, X. Li, J. Yang, H. Wu, J. Chen, W. Hu, X. Wang, Y. Xu, Z. Wang, Y. Xu, J. Wang, D. Sahoo, T. Yu, and C. Xiong · 2025
Closest in time.
Dream 7b, 2025
J. Ye, Z. Xie, L. Zheng, J. Gao, Z. Wu, X. Jiang, Z. Li, and L. Kong · 2025
Closest in time.