Fetching the paper…
Reading the bibliography…
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models.
Understanding the difficulty of training transformers
L. Liu, X. Liu, J. Gao, W. Chen, and J. Han · 2004
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Large scale distributed neural network training through online distillation, 2018
R. Anil, G. Pereyra, A. Passos, R. Ormandi, G. E. Dahl, and G. E. Hinton · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, C. Xu, and J. J. Corso · 2018
Earlier work this paper cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang · 2019
Earlier work this paper cites.
ActivityNet-QA: A dataset for understanding complex web videos via question answering
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao · 2019
Earlier work this paper cites.
GShard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2020
Earlier work this paper cites.
Covost 2: A massively multilingual speech-to-text translation corpus, 2020
C. Wang, A. Wu, and J. Pino · 2020
Earlier work this paper cites.
GLaM: Efficient scaling of language models with mixture-of-experts
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, et al · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2021
Earlier work this paper cites.
Detecting moments and highlights in videos via natural language queries
J. Lei, T. L. Berg, and M. Bansal · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts, 2021
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. S. Pinto, D. Keysers, and N. Houlsby · 2021
Earlier work this paper cites.
Hash layers for large sparse models
S. Roller, S. Sukhbaatar, J. Weston, et al · 2021
Earlier work this paper cites.
Mlp-mixer: An all-mlp architecture for vision, 2021
I. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit, M. Lucic, and A. Dosovitskiy · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, et al · 2022
Earlier work this paper cites.
Pathways: Asynchronous distributed dataflow for ml
P. Barham, A. Chowdhery, J. Dean, S. Ghemawat, S. Hand, D. Hurt, M. Isard, H. Lim, R. Pang, S. Roy, et al · 2022
Earlier work this paper cites.
Quantifying memorization across neural language models
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang · 2022
Earlier work this paper cites.
PaLM: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Earlier work this paper cites.
Unified scaling laws for routed language models, 2022
A. Clark, D. de las Casas, A. Guy, A. Mensch, M. Paganini, J. Hoffmann, B. Damoc, B. Hechtman, T. Cai, S. Borgeaud, G. van den Driessche, E. Rutherford, T. Hennigan, M. Johnson, K. Millican, A. Cassirer, C. Jones, E. Buchatskaya, D. Budden, L. Sifre, S. Osindero, O. Vinyals, J. Rae, E. Elsen, K. Kavukcuoglu, and K. Simonyan · 2022
Earlier work this paper cites.
Preventing verbatim memorization in language models gives a false sense of privacy, 2022
D. Ippolito, F. Tramer, M. Nasr, C. Zhang, M. Jagielski, K. Lee, C. A. Choquette-Choo, and N. Carlini · 2022
Earlier work this paper cites.
Red teaming language models with language models, 2022
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Earlier work this paper cites.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, et al · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, et al · 2023
Earlier work this paper cites.
Fleurs: Few-shot learning evaluation of universal representations of speech
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna · 2023
Earlier work this paper cites.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al · 2023
Earlier work this paper cites.
Vectara Hallucination Leaderboard, nov 2023
S. Hughes, M. Bae, and M. Li · 2023
Earlier work this paper cites.
MADLAD-400: A Multilingual And Document-Level Large Audited Dataset, 2023
S. Kudugunta, I. Caswell, B. Zhang, X. Garcia, C. A. Choquette-Choo, K. Lee, D. Xin, A. Kusupati, R. Stella, A. Bapna, and O. Firat · 2023
Earlier work this paper cites.
A theory on adam instability in large-scale machine learning
I. Molybog, P. Albert, M. Chen, Z. DeVito, D. Esiobu, N. Goyal, P. Koura, S. Narang, A. Poulton, R. Silva, et al · 2023
Earlier work this paper cites.
Scalable extraction of training data from (production) language models, 2023
M. Nasr, N. Carlini, J. Hayase, M. Jagielski, A. F. Cooper, D. Ippolito, C. A. Choquette-Choo, E. Wallace, F. Tramèr, and K. Lee · 2023
Cited alongside, same era.
Perception test: A diagnostic benchmark for multimodal video models
V. Patraucean, L. Smaira, A. Gupta, A. Recasens, L. Markeeva, D. Banarse, S. Koppula, M. Malinowski, Y. Yang, C. Doersch, et al · 2023
Cited alongside, same era.
Small-scale proxies for large-scale transformer training instabilities
M. Wortsman, P. J. Liu, L. Xiao, K. Everett, A. Alemi, B. Adlam, J. D. Co-Reyes, I. Gur, A. Kumar, R. Novak, et al · 2023
Cited alongside, same era.
Intercode: Standardizing and benchmarking interactive coding with execution feedback, 2023
J. Yang, A. Prabhakar, K. Narasimhan, and S. Yao · 2023
Cited alongside, same era.
Stabilizing transformer training by preventing attention entropy collapse
Holistic safety and responsibility evaluations of advanced ai models, 2024
L. Weidinger, J. Barnhart, J. Brennan, C. Butterfield, S. Young, W. Hawkins, et al · 2024
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2024
Later among the works it cites.
Pokemon Red Version - Guide and Walkthrough (GB), 2024
Zerokid · 2024
Later among the works it cites.
Claude’s extended thinking, 2025
Anthropic · 2025
Closest in time.
Advancing the frontier of video understanding with Gemini 2.5, 2025
A. Baddepudi, A. Yang, and M. Lučić · 2025
Closest in time.
Matharena: Evaluating llms on uncontaminated math competitions, 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Zhai, T. Likhomanenko, E. Littwin, D. Busbridge, J. Ramapuram, Y. Zhang, J. Gu, and J. M. Susskind · 2023
Cited alongside, same era.
A. Beutel, K. Xiao, J. Heidecke, and L. Weng · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. Jordan, J. E. Gonzalez, et al · 2024
Cited alongside, same era.
Introducing SWE-bench verified, 2024
N. Chowdhury, J. Aung, C. J. Shern, O. Jaffe, D. Sherburn, G. Starace, E. Mays, R. Dias, M. Aljubeh, M. Glaese, C. E. Jimenez, J. Yang, L. Ho, T. Patwardhan, K. Liu, and A. Madry · 2024
Cited alongside, same era.
CodeGemma: Open Code Models Based on Gemma, 2024
CodeGemma Team, H. Zhao, J. Hui, J. Howland, N. Nguyen, S. Zuo, A. Hu, C. A. Choquette-Choo, J. Shen, J. Kelley, K. Bansal, L. Vilnis, M. Wirth, P. Michel, P. Choy, P. Joshi, R. Kumar, S. Hashmi, S. Agrawal, Z. Gong, J. Fine, T. Warkentin, A. J. Hartman, B. Ni, K. Korevec, K. Schaefer, and S. Huffman · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team · 2024
Cited alongside, same era.
Gemini Deep Research, 2024
Gemini Team, Google · 2024
Cited alongside, same era.
Gemma: Open Models Based on Gemini Research and Technology, 2024
Gemma Team · 2024
Cited alongside, same era.
M. Balunović, J. Dekoninck, I. Petrov, N. Jovanović, and M. Vechev · 2025
Closest in time.
Gemini 2.5: Our most intelligent models are getting even better, 2025b
T. Doshi · 2025
Closest in time.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang, et al · 2025
Closest in time.
Aider Polyglot Coding Leaderboard, 2025
P. Gauthier · 2025
Closest in time.
Eclektic: a novel challenge set for evaluation of cross-lingual knowledge transfer, 2025
O. Goldman, U. Shaham, D. Malkin, S. Eiger, A. Hassidim, Y. Matias, J. Maynez, A. M. Gilady, J. Riesa, S. Rijhwani, L. Rimell, I. Szpektor, R. Tsarfaty, and M. Eyal · 2025
Closest in time.
Our vision for building a universal AI assistant, 2025
D. Hassabis · 2025
Closest in time.
Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos, 2025
K. Hu, P. Wu, F. Pu, W. Xiao, Y. Zhang, X. Yue, B. Li, and Z. Liu · 2025
Closest in time.
The facts grounding leaderboard: Benchmarking llms’ ability to ground responses to long-form input
A. Jacovi, A. Wang, C. Alberti, C. Tao, J. Lipovetz, K. Olszewska, L. Haas, M. Liu, N. Keating, A. Bloniarz, et al · 2025
Closest in time.
Experiment with Gemini 2.0 Flash native image generation, 2025
K. Kampf and N. Brichtova · 2025
Closest in time.
Gemini 2.0 is now available to everyone, 2025
K. Kavukcuoglu · 2025
Closest in time.
Gemini 2.5 Pro Preview: even better coding performance, 2025
L. Kilpatrick · 2025
Closest in time.
Evaluating Gemini in an Arena for Learning, 2025
LearnLM Team · 2025
Closest in time.
Webdev arena, 2025
LMArena Team · 2025
Closest in time.
Gemini 2.0: Flash, Flash-Lite and Pro, 2025
S. B. Mallick and L. Kilpatrick · 2025
Closest in time.
L. Phan et al · 2025
Closest in time.
Evaluating frontier models for stealth and situational awareness, 2025
M. Phuong, R. S. Zimmermann, Z. Wang, D. Lindner, V. Krakovna, S. Cogan, A. Dafoe, L. Ho, and R. Shah · 2025
Closest in time.
Google I/O 2025: From research to reality, 2025
S. Pichai · 2025
Closest in time.
Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos
C. Plizzari, A. Tonioni, Y. Xian, A. Kulshrestha, and F. Tombari · 2025
Closest in time.
ZeroBench: An impossible visual benchmark for contemporary large multimodal models
J. Roberts, M. R. Taesiri, A. Sharma, A. Gupta, S. Roberts, I. Croitoru, S.-V. Bogolin, J. Tang, F. Langer, V. Raina, et al · 2025
Closest in time.
A framework for evaluating emerging cyberattack capabilities of AI, 2025
M. Rodriguez, R. A. Popa, L. Liang, A. Wang, M. Rahtz, A. Kaskasoli, A. Dafoe, and F. Flynn · 2025
Closest in time.
An approach to technical agi safety and security, 2025
R. Shah, A. Irpan, A. M. Turner, A. Wang, A. Conmy, D. Lindner, J. Brown-Cohen, L. Ho, N. Nanda, R. A. Popa, R. Jain, R. Greig, S. Albanie, S. Emmons, S. Farquhar, S. Krier, S. Rajamanoharan, S. Bridgers, T. Ijitoye, T. Everitt, V. Krakovna, V. Varma, V. Mikulik, Z. Kenton, D. Orr, S. Legg, N. Goodman, A. Dafoe, F. Flynn, and A. Dragan · 2025
Closest in time.
Upload and edit your images directly in the Gemini app, 2025
D. Sharon · 2025
Closest in time.
Lessons from defending gemini against indirect prompt injections, 2025
C. Shi, S. Lin, S. Song, J. Hayes, I. Shumailov, I. Yona, J. Pluto, A. Pappu, C. A. Choquette-Choo, M. Nasr, C. Sitawarin, G. Gibson, A. Terzis, and J. F. Flynn · 2025
Closest in time.
Expanding AI Overviews and introducing AI Mode, 2025
R. Stein · 2025
Closest in time.
H. Wijk, T. Lin, J. Becker, S. Jawhar, N. Parikh, T. Broadley, L. Chan, M. Chen, J. Clymer, J. Dhyani, et al · 2025
Closest in time.
Gemini Plays Pokemon Twitch Stream, 2025
J. Zhang · 2025
Closest in time.