Fetching the paper…
Reading the bibliography…
The emergence of deep research systems presents significant capabilities in problem-solving, extending from basic queries to sophisticated research tasks.
The use of multiple measurements in taxonomic problems
Ronald A. Fisher. 1936 · 1936
Earlier work this paper cites.
The measurement of observer agreement for categorical data
J Richard Landis and Gary G Koch. 1977 · 1977
Earlier work this paper cites.
Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation
David M. W. Powers. 2011 · 2011
Earlier work this paper cites.
Atypical combinations and scientific impact
Brian Uzzi, Satyam Mukherjee, Michael Stringer, and Ben Jones. 2013 · 2013
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Citations, citation indicators, and research quality: An overview of basic concepts and theories
Dag W Aksnes, Liv Langfeldt, and Paul Wouters. 2019 · 2019
Earlier work this paper cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023 · 2023
Earlier work this paper cites.
Openscholar: Synthesizing scientific literature with retrieval-augmented lms
Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D’arcy, et al. 2024 · 2024
Earlier work this paper cites.
Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data
Florian E Dorner, Vivian Y Nastl, and Moritz Hardt. 2024 · 2024
Earlier work this paper cites.
Ragas: Supercharge your llm application evaluations
ExplodingGradients. 2024 · 2024
Earlier work this paper cites.
Sciknoweval: Evaluating multi-level scientific knowledge of large language models
Kehua Feng, Keyan Ding, Weijie Wang, Xiang Zhuang, Zeyuan Wang, Ming Qin, Yu Zhao, Jianhua Yao, Qiang Zhang, and Huajun Chen. 2024 · 2024
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024 · 2024
Earlier work this paper cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024 · 2024
Earlier work this paper cites.
The ai scientist: Towards fully automated open-ended scientific discovery
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024 · 2024
Earlier work this paper cites.
Introducing chatgpt search
OpenAI. 2024 · 2024
Earlier work this paper cites.
Towards human-ai mutual learning: A new research paradigm
Xiaomei Wang and Xiaoyu Chen. 2024 · 2024
Cited alongside, same era.
Corag: A cost-constrained retrieval optimization system for retrieval-augmented generation
Ziting Wang, Haitao Yuan, Wei Dong, Gao Cong, and Feifei Li. 2024 · 2024
Cited alongside, same era.
Openresearcher: Unleashing ai for accelerated scientific research
Yuxiang Zheng, Shichao Sun, Lin Qiu, Dongyu Ru, Cheng Jiayang, Xuefeng Li, Jifan Lin, Binjie Wang, Yun Luo, Renjie Pan, et al. 2024 · 2024
Cited alongside, same era.
Open deep search: Democratizing search with open-source reasoning agents
Salaheddin Alzubi, Creston Brooks, Purva Chiniya, Edoardo Contente, Chiara von Gerlach, Lucas Irwin, Yihan Jiang, Arda Kaz, Windsor Nguyen, Sewoong Oh, et al. 2025 · 2025
Cited alongside, same era.
Swe-lancer: Can frontier llms earn $1 million from real-world freelance software engineering?
Samuel Miserendino, Michele Wang, Tejal Patwardhan, and Johannes Heidecke. 2025 · 2025
Closest in time.
Deep research system card
OpenAI. 2025 · 2025
Closest in time.
Gpt-4o search preview
OpenAI. 2025 · 2025
Closest in time.
Introducing deep research
OpenAI. 2025 · 2025
Closest in time.
Introducing perplexity deep research
Perplexity AI. 2025 · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Claude 3.7 sonnet system card
Anthropic. 2025 · 2025
Cited alongside, same era.
Learning to reason with search for llms via reinforcement learning
Mingyang Chen, Tianpeng Li, Haoze Sun, Yijie Zhou, Chenzheng Zhu, Fan Yang, Zenan Zhou, Weipeng Chen, Haofen Wang, Jeff Z Pan, et al. 2025 · 2025
Cited alongside, same era.
Gemini 2.5 flash
Google DeepMind. 2025a · 2025
Cited alongside, same era.
Gemini 2.5 pro
Google DeepMind. 2025b · 2025
Cited alongside, same era.
Deepresearch bench: A comprehensive benchmark for deep research agents
Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. 2025 · 2025
Cited alongside, same era.
Gemini deep research - your personal research assistant
Google. 2025 · 2025
Cited alongside, same era.
Mind2web 2: Evaluating agentic search with agent-as-a-judge
Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu, Michael Lin, Weijian Qi, Andrei Kopanev, Botao Yu, Bernal Jiménez Gutiérrez, Yiheng Shu, et al. 2025 · 2025
Cited alongside, same era.
Agentic ai for scientific discovery: A survey of progress, challenges, and future directions
Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack. 2025 · 2025
Cited alongside, same era.
Accelerating scientific breakthroughs with an ai co-scientist
Google Research. 2025 · 2025
Closest in time.
R1-searcher: Incentivizing the search capability in llms via reinforcement learning
Huatong Song, Jinhao Jiang, Yingqian Min, Jie Chen, Zhipeng Chen, Wayne Xin Zhao, Lei Fang, and Ji-Rong Wen. 2025 · 2025
Closest in time.
The 2025 ai index report
Stanford HAI. 2025 · 2025
Closest in time.
Paperbench: Evaluating ai’s ability to replicate ai research
Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, et al. 2025 · 2025
Closest in time.
Browsecomp: A simple yet challenging benchmark for browsing agents
Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. 2025 · 2025
Closest in time.
From algorithms to academia: An endeavor to benchmark ai-generated scientific papers against human standards
Jackson Woodrow, Nour Nassour, John Y Kwon, Soheil Ashkani-Esfahani, and Mitchel Harris. 2025 · 2025
Closest in time.
Grok 3 beta — the age of reasoning agents
xAI. 2025 · 2025
Closest in time.
A comprehensive survey of deep research: Systems, methodologies, and applications
Renjun Xu and Jingwen Peng. 2025 · 2025
Closest in time.
Deepresearcher: Scaling deep research via reinforcement learning in real-world environments
Yuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, and Pengfei Liu. 2025 · 2025
Closest in time.