Fetching the paper…
Reading the bibliography…
Information tasks such as writing surveys or analytical reports require complex search and reasoning, and have recently been grouped under the umbrella of \textit{deep research} -- a term also adopted by recent models targeting these capabilities.
Patent claims revisited
Dargaye Churnet · 2012
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Eli5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 2019
Earlier work this paper cites.
Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa · 2020
Earlier work this paper cites.
Cuad: An expert-annotated nlp dataset for legal contract review
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball · 2021
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants
Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2023
Earlier work this paper cites.
Openscholar: Synthesizing scientific literature with retrieval-augmented lms
Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D’arcy, et al · 2024
Earlier work this paper cites.
Claim verification in the age of large language models: A survey
Alphaeus Dmonte, Roland Oruche, Marcos Zampieri, Prasad Calyam, and Isabelle Augenstein · 2024
Earlier work this paper cites.
Analysis of plan-based retrieval for grounded text generation
Ameya Godbole, Nicholas Monath, Seungyeon Kim, Ankit Singh Rawat, Andrew McCallum, and Manzil Zaheer · 2024
Earlier work this paper cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Earlier work this paper cites.
Researcharena: Benchmarking llms’ ability to collect and organize information as research agents
Hao Kang and Chenyan Xiong · 2024
Earlier work this paper cites.
Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation, 2024
Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, and Manaal Faruqui · 2024
Earlier work this paper cites.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2024
Earlier work this paper cites.
Assisting in writing wikipedia-like articles from scratch with large language models
Yijia Shao, Yucheng Jiang, Theodore A Kanell, Peter Xu, Omar Khattab, and Monica S Lam · 2024
Earlier work this paper cites.
Deerflow: Deep exploration and efficient research flow
Henry Li Bytedance, Daniel Walnut · 2025
Cited alongside, same era.
Deep research comparator: A platform for fine-grained human annotations of deep research agents
Prahaladh Chandrahasan, Jiahe Jin, Zhihan Zhang, Tevin Wang, Andy Tang, Lucy Mo, Morteza Ziyadi, Leonardo FR Ribeiro, Zimeng Qiu, Markus Dreyer, et al · 2025
Cited alongside, same era.
Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research, 2025
João Coelho, Jingjie Ning, Jingyuan He, Kangrui Mao, Abhijay Paladugu, Pranav Setlur, Jiahe Jin, Jamie Callan, João Magalhães, Bruno Martins, and Chenyan Xiong · 2025
Cited alongside, same era.
Deepresearch bench: A comprehensive benchmark for deep research agents, 2025
Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao · 2025
Cited alongside, same era.
Deep research bench: Evaluating ai web research agents, 2025
Introducing openai o3 and o4-mini
OpenAI · 2025
Closest in time.
Introducing deep research
OpenAI · 2025
Closest in time.
Perplexity deep research
PerplexityAI · 2025
Closest in time.
Sonar pro
PerplexityAI · 2025
Closest in time.
Sonar reasoning
PerplexityAI · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al · 2025
Closest in time.
Pangu deepdiver: Adaptive search intensity scaling via open-web reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
FutureSearch, :, Nikos I. Bosse, Jon Evans, Robert G. Gambee, Daniel Hnyk, Peter Mühlbacher, Lawrence Phillips, Dan Schwarz, and Jack Wildman · 2025
Cited alongside, same era.
We’re expanding our gemini 2.5 family of models
Google · 2025
Cited alongside, same era.
Deep research is now available on gemini 2.5 pro experimental
Google · 2025
Cited alongside, same era.
Precise information control in long-form text generation
Jacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li, Zhiyuan Zeng, Weijia Shi, Yulia Tsvetkov, Danqi Chen, Pang Wei Koh, and Luke Zettlemoyer · 2025
Cited alongside, same era.
Deep research agents: A systematic examination and roadmap, 2025
Yuxuan Huang, Yihang Chen, Haozheng Zhang, Kang Li, Meng Fang, Linyi Yang, Xiaoguang Li, Lifeng Shang, Songcen Xu, Jianye Hao, Kun Shao, and Jun Wang · 2025
Cited alongside, same era.
Webthinker: Empowering large reasoning models with deep research capability
Xiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian, Yutao Zhu, Yongkang Wu, Ji-Rong Wen, and Zhicheng Dou · 2025
Cited alongside, same era.
Veritrail: Closed-domain hallucination detection with traceability
Dasha Metropolitansky and Jonathan Larson · 2025
Cited alongside, same era.
Introducing researcher and analyst in microsoft 365 copilot
Microsoft and Jared Spataro · 2025
Cited alongside, same era.
Wenxuan Shi, Haochen Tan, Chuqiao Kuang, Xiaoguang Li, Xiaozhe Ren, Chen Zhang, Hanting Chen, Yasheng Wang, Lifeng Shang, Fisher Yu, et al · 2025
Closest in time.
Geak: Introducing triton kernel ai agent & evaluation benchmarks
Jianghui Wang, Vinay Joshi, Saptarshi Majumder, Xu Chao, Bin Ding, Ziqiong Liu, Pratik Prabhanjan Brahma, Dong Li, Zicheng Liu, and Emad Barsoum · 2025
Closest in time.
Grok 3 beta — the age of reasoning agents
xAI · 2025
Closest in time.
A comprehensive survey of deep research: Systems, methodologies, and applications, 2025
Renjun Xu and Jingwen Peng · 2025
Closest in time.
Researcherbench: Evaluating deep ai research systems on the frontiers of scientific inquiry
Tianze Xu, Pengrui Lu, Lyumanshan Ye, Xiangkun Hu, and Pengfei Liu · 2025
Closest in time.
Open deep research
David Zhang · 2025
Closest in time.
Deepresearcher: Scaling deep research via reinforcement learning in real-world environments
Yuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, and Pengfei Liu · 2025
Closest in time.