Fetching the paper…
Reading the bibliography…
Data contamination refers to the leakage of evaluation data into model training data, resulting in overfitting to supposedly held-out test sets and compromising test validity.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2023
Earlier work this paper cites.
C. Deng, Y. Zhao, Y. Heng, Y. Li, J. Cao, X. Tang, and A. Cohan · 2024
Earlier work this paper cites.
Y. Dong, X. Jiang, H. Liu, Z. Jin, B. Gu, M. Yang, and G. Li · 2024
Earlier work this paper cites.
Me, myself, and ai: The situational awareness dataset (sad) for llms
R. Laine, B. Chughtai, J. Betley, K. Hariharan, M. Balesni, J. Scheurer, M. Hobbhahn, A. Meinke, and O. Evans · 2024
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2024
Earlier work this paper cites.
Measuring short-form factuality in large language models
J. Wei, N. Karina, H. W. Chung, Y. J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus · 2024
Cited alongside, same era.
A careful examination of large language model performance on grade school arithmetic
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, C. Zhuang, D. Slack, et al · 2024
Cited alongside, same era.
Deep research comparator: A platform for fine-grained human annotations of deep research agents
P. Chandrahasan, J. Jin, Z. Zhang, T. Wang, A. Tang, L. Mo, M. Ziyadi, L. F. R. Ribeiro, Z. Qiu, M. Dreyer, A. Asai, and C. Xiong · 2025
Cited alongside, same era.
Deepresearch bench: A comprehensive benchmark for deep research agents
M. Du, B. Xu, C. Zhu, X. Wang, and Z. Mao · 2025
Cited alongside, same era.
Mind2web 2: Evaluating agentic search with agent-as-a-judge, 2025
B. Gou, Z. Huang, Y. Ning, Y. Gu, M. Lin, W. Qi, A. Kopanev, B. Yu, B. J. Gutiérrez, Y. Shu, C. H. Song, J. Wu, S. Chen, H. N. Moussa, T. Zhang, J. Xie, Y. Li, T. Xue, Z. Liao, K. Zhang, B. Zheng, Z. Cai, V. Rozgic, M. Ziyadi, H. Sun, and Y. Su · 2025
Search arena: Analyzing search-augmented llms
M. Miroyan, T.-H. Wu, L. King, T. Li, J. Pan, X. Hu, W.-L. Chiang, A. N. Angelopoulos, T. Darrell, N. Norouzi, and J. Gonzalez · 2025
Closest in time.
L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al · 2025
Closest in time.
Bearcubs: A benchmark for computer-using web agents
Y. Song, K. Thai, C. M. Pham, Y. Chang, M. Nadaf, and M. Iyyer · 2025
Closest in time.
Browsecomp: A simple yet challenging benchmark for browsing agents
J. Wei, Z. Sun, S. Papay, S. McKinney, J. Han, I. Fulford, H. W. Chung, A. T. Passos, W. Fedus, and A. Glaese · 2025
Closest in time.
Infodeepseek: Benchmarking agentic information seeking for retrieval-augmented generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gemini deep research - your personal research assistant
G. DeepMind
Cited in the paper.
URL https://openai.com/index/introducing-deep-research/
OpenAI, a
Cited in the paper.
URL https://cdn.openai.com/pdf/8124a3ce-ab78-4f06-96eb-49ea29ffb52f/gpt5-system-card-aug7.pdf
OpenAI, b
Cited in the paper.
URL https://www.perplexity.ai/hub/blog/introducing-perplexity-deep-research
Perplexity
Cited in the paper.
URL https://x.ai/news/grok-4
XAI
Cited in the paper.
Y. Xi, J. Lin, M. Zhu, Y. Xiao, Z. Ou, J. Liu, T. Wan, B. Chen, W. Liu, Y. Wang, R. Tang, W. Zhang, and Y. Yu · 2025
Closest in time.