Fetching the paper…
Reading the bibliography…
Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R · 2017
Earlier work this paper cites.
On the measure of intelligence
Chollet, F · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
”Transformers: State-of-the-Art Natural Language Processing”
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Earlier work this paper cites.
Scaling scaling laws with board games
Jones, A. L · 2021
Earlier work this paper cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Northcutt, C. G., Athalye, A., and Mueller, J · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
A Benchmark for Compositional Visual Reasoning
Zerroug, A., Vaishnav, M., Colin, J., Musslick, S., and Serre, T · 2022
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Earlier work this paper cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Seed-Bench: Benchmarking multimodal llms with generative comprehension
Li, B., Wang, R., Wang, G., Ge, Y., Ge, Y., and Shan, Y · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2023
Earlier work this paper cites.
Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
Mensink, T., Uijlings, J., Castrejon, L., Goel, A., Cadar, F., Zhou, H., Sha, F., Araujo, A., and Ferrari, V · 2023
Earlier work this paper cites.
OpenCompass: A Universal Evaluation Platform for Foundation Models
OpenCompass Contributors · 2023
Earlier work this paper cites.
GPQA: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Earlier work this paper cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L · 2023
Earlier work this paper cites.
MM-Vet: Evaluating large multimodal models for integrated capabilities
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L · 2023
Earlier work this paper cites.
Scaling Test Time Compute with Open Models
Beeching, E., Tunstall, L., and Rush, S · 2024
Earlier work this paper cites.
OpenAI o3 Breakthrough High Score on ARC-AGI-Pub
Chollet, F · 2024
Cited alongside, same era.
ARC prize 2024: Technical report
Chollet, F., Knoop, M., Kamradt, G., and Landers, B · 2024
Cited alongside, same era.
NVLM: Open frontier-class multimodal llms
Dai, W., Lee, N., Wang, B., Yang, Z., Liu, Z., Barker, J., Rintamaki, T., Shoeybi, M., Catanzaro, B., and Ping, W · 2024
Cited alongside, same era.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al · 2024
Cited alongside, same era.
ReMI: A Dataset for Reasoning with Multiple Images
Kazemi, M., Dikkala, N., Anand, A., Devic, P., Dasgupta, I., Liu, F., Fatemi, B., Awasthi, P., Guo, D., Gollapudi, S., et al · 2024
Humaneval-v: Benchmarking high-level visual reasoning with complex diagrams in coding tasks
Zhang, F., Wu, L., Bai, H., Lin, G., Li, X., Yu, X., Wang, Y., Chen, B., and Keung, J · 2024
Later among the works it cites.
AI/ML API Inference Pricing
AI/ML API · 2025
Closest in time.
Blink: Multimodal large language models can see but not perceive
Fu, X., Hu, Y., Li, B., Feng, Y., Wang, H., Lin, X., Roth, D., Smith, N. A., Ma, W.-C., and Krishna, R · 2025
Closest in time.
Access the latest 2.0 experimental models in the gemini app
Google · 2025
Closest in time.
Gemini 2.0 flash model card
Google DeepMind · 2025
Closest in time.
LLMTest_NeedleInAHaystack
Kamradt, G · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey on benchmarks of multimodal large language models
Li, J., Lu, W., Fei, H., Luo, M., Dai, M., Xia, M., Jin, Y., Gan, Z., Qi, D., Fu, C., et al · 2024
Cited alongside, same era.
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Miller, E · 2024
Cited alongside, same era.
Pixtral large
Mistral AI · 2024
Cited alongside, same era.
LHRS-Bot: Empowering remote sensing with vgi-enhanced large multimodal language model
Muhtar, D., Li, Z., Gu, F., Zhang, X., and Xiao, P · 2024
Cited alongside, same era.
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
Padlewski, P., Bain, M., Henderson, M., Zhu, Z., Relan, N., Pham, H., Ong, D., Aleksiev, K., Ormazabal, A., Phua, S., et al · 2024
Cited alongside, same era.
QVQ: To See the World with Wisdom, 2024
Qwen-Team · 2024
Cited alongside, same era.
Vision language models are blind
Rahmanzadehgervi, P., Bolton, L., Taesiri, M. R., and Nguyen, A. T · 2024
Cited alongside, same era.
Gemini 2.5: Our most intelligent AI model
Kavukcuoglu, K · 2025
Closest in time.
MMBench: Is your multi-modal model an all-around player?
Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al · 2025
Closest in time.
FrontierCS: Evolving challenges for evolving intelligence
Mang, Q., Chai, W., Li, Z., Mao, H., Zhou, S., Du, A., Li, H., Liu, S., Chen, E., Wang, Y., Chu, X., Cheng, Z., Xu, Y., Xia, T., Wang, Z., Shi, T., Yao, J., Zhao, Y., Zhang, Q., Ruan, C., Shen, Z., Liu, K., He, R., Xing, D., Li, Z., Zeng, Z., Jiang, Y., Cheng, L., Zhao, Z., Sun, Y., Zheng, W., Zhang, M., Ji, R., Tu, X., Zheng, Z., Chen, Z., Zhou, K., Wang, Z., Chen, J., Korolova, A., Henderson, P., Viswanath, P., Ganesh, V., Xie, S., Liu, Z., Song, D., Min, S., Stoica, I., Gonzalez, J. E., Shang, J., and Cheung, A · 2025
Closest in time.
The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation
Meta · 2025
Closest in time.
Mistral AI API (0.0.2)
Mistral · 2025
Closest in time.
GDPval: Evaluating AI model performance on real-world economically valuable tasks
Patwardhan, T., Dias, R., Proehl, E., Kim, G., Wang, M., Watkins, O., Posada Fishman, S., Aljubeh, M., Thacker, P., Fauconnet, L., Kim, N. S., Chao, P., Miserendino, S., Chabot, G., Li, D., Sharman, M., Barr, A., Glaese, A., and Tworek, J · 2025
Closest in time.
Reka AI API
Reka · 2025
Closest in time.
GRAB: A challenging graph analysis benchmark for large multimodal models
Roberts, J., Han, K., and Albanie, S · 2025
Closest in time.
Agent Laboratory: Using LLM Agents as Research Assistants
Schmidgall, S., Su, Y., Wang, Z., Sun, X., Wu, J., Yu, X., Liu, J., Liu, Z., and Barsoum, E · 2025
Closest in time.
FrontierScience: Evaluating AI’s ability to perform expert-level scientific tasks, 2025
Wang, M., Lin, R., Hu, K., Jiao, J., Chowdhury, N., Chang, E., and Patwardhan, T · 2025
Closest in time.
Grok 4 model card
xAI · 2025
Closest in time.
ARC-AGI leaderboard
ARC Prize Foundation · 2026
Closest in time.
Gemini API release notes (changelog)
Google · 2026
Closest in time.
Gemini 3.1 Pro model card
Google DeepMind · 2026
Closest in time.