Fetching the paper…
Reading the bibliography…
We present BrowseComp, a simple yet challenging benchmark for measuring the ability for agents to browse the web.
Building Watson: An overview of the DeepQA project
D. Ferrucci, E. Brown, J. Chu-Carroll, J. Fan, D. Gondek, A. Kalyanpur, A. Lally, J. W. Murdock, E. Nyberg, J. Prager, N. Schlaefer, and C. Welty · 2010
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and verification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
E. Dinan, S. Roller, K. Shuster, A. Fan, M. Auli, and J. Weston · 2019
Earlier work this paper cites.
ELI5: Long form question answering
A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
K. Lee, M.-W. Chang, and K. Toutanova · 2019
Earlier work this paper cites.
REALM: Retrieval-augmented language model pre-training
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. tau Yih, T. Rocktäschel, S. Riedel, and D. Kiela · 2020
Cited alongside, same era.
WebGPT: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman · 2021
Cited alongside, same era.
KILT: a benchmark for knowledge intensive language tasks
F. Petroni, A. Piktus, A. Fan, P. Lewis, M. Yazdani, N. De Cao, J. Thorne, Y. Jernite, V. Karpukhin, J. Maillard, V. Plachouras, T. Rocktäschel, and S. Riedel · 2021
Cited alongside, same era.
Teaching language models to support answers with verified quotes
J. Menick, T. Miller, T. Ribeiro, L. G. Brigato, A. Bakhtin, P. Chatelain, M. Moore, J. Kramár, A. Joulin, K. Shuster, P. Stenetorp, and D. Kiela · 2022
Cited alongside, same era.
Webshop: Towards scalable real-world web interaction with grounded language agents
S. Yao, H. Chen, J. Yang, and K. Narasimhan · 2022
Try Deep Research and our new experimental model in Gemini, your AI assistant
Google · 2024
Later among the works it cites.
GPQA: A graduate-level google-proof Q&A benchmark
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2024
Later among the works it cites.
Measuring short-form factuality in large language models
J. Wei, N. Karina, H. W. Chung, Y. J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus · 2024
Later among the works it cites.
Introducing perplexity deep research
perplexity.AI · 2025
Closest in time.
Humanity’s last exam
L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al · 2025
Closest in time.
BEARCUBS: A benchmark for computer-using web agents
Y. Song, K. Thai, C. M. Pham, Y. Chang, M. Nadaf, and M. Iyyer · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GAIA: a benchmark for general AI assistants
G. Mialon, C. Fourrier, T. Wolf, Y. LeCun, and T. Scialom · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolatingthe capabilities of language models
A. Srivastava, D. Kleyko, and Z. Wu · 2023
Cited alongside, same era.
Introducing deep research
OpenAI
Cited in the paper.
Introducing operator
OpenAI
Cited in the paper.
Grok 3 beta — the age of reasoning agents
x.AI · 2025
Closest in time.
ReAct: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2025
Closest in time.