Fetching the paper…
Reading the bibliography…
As large language models (LLMs) evolve into tool-using agents, the ability to browse the web in real-time has become a critical yardstick for measuring their reasoning and retrieval competence.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Retrieval augmented language model pre-training. In International conference on machine learning . PMLR, 3929–3938
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
KILT: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, et al · 2020
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
The web can be your oyster for improving large language models
Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jingyuan Wang, Jian-Yun Nie, and Ji-Rong Wen. 2023a · 2023
Earlier work this paper cites.
Huatuo-26m, a large-scale chinese medical qa dataset
Jianquan Li, Xidong Wang, Xiangbo Wu, Zhiyi Zhang, Xiaolong Xu, Jie Fu, Prayag Tiwari, Xiang Wan, and Benyou Wang. 2023b · 2023
Earlier work this paper cites.
Benchmarking large language models on cmexam-a comprehensive chinese medical exam dataset
Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Andrew Liu, Helin Wang, Chenyu You, Zhenhua Guo, Lei Zhu, et al · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants. In The Twelfth International Conference on Learning Representations
Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023 · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Earlier work this paper cites.
Freshllms: Refreshing large language models with search engine augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, et al · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR)
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023 · 2023
Earlier work this paper cites.
Introducing Claude 3.5 Sonnet
Anthropic. 2024 · 2024
Earlier work this paper cites.
From local to global: A graph rag approach to query-focused summarization
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024 · 2024
Cited alongside, same era.
A survey on rag meeting llms: Towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 6491–6501
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024 · 2024
Cited alongside, same era.
Marcos Fernández-Pichel, Juan C Pichel, and David E Losada. 2024 · 2024
Cited alongside, same era.
Introducing Gemini 2.0: our new AI model for the agentic era
Google. 2024b · 2024
Cited alongside, same era.
Zero-Indexing Internet Search Augmented Generation for Large Language Models
Introducing Claude 3.7 Sonnet and Claude Code
Anthropic. 2025 · 2025
Closest in time.
https://www.doubao.com/chat/
ByteDance. 2025 · 2025
Closest in time.
https://chat.deepseek.com/
DeepSeek. 2025 · 2025
Closest in time.
Gemini 2.5: Our most intelligent AI model
Google. 2024a · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
WebGLM: Towards an Efficient and Reliable Web-Enhanced Question Answering System
Hanyu Lai, Xiao Liu, Hao Yu, Yifan Xu, Iat Long Iong, Shuntian Yao, Aohan Zeng, Zhengxiao Du, Yuxiao Dong, and Jie Tang. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guangxin He, Zonghong Dai, Jiangcheng Zhu, Binqiang Zhao, Qicheng Hu, Chenyue Li, You Peng, Chen Wang, and Binhang Yuan. 2024 · 2024
Cited alongside, same era.
Level-Navi Agent: A Framework and benchmark for Chinese Web Search Agents
Chuanrui Hu, Shichong Xie, Baoxin Wang, Bin Chen, Xiaofeng Cong, and Jun Zhang. 2024 · 2024
Cited alongside, same era.
AutoWebGLM: A Large Language Model-based Web Navigating Agent. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 5295–5306
Hanyu Lai, Xiao Liu, Iat Long Iong, Shuntian Yao, Yuxuan Chen, Pengbo Shen, Hao Yu, Hanchen Zhang, Xiaohan Zhang, Yuxiao Dong, et al · 2024
Cited alongside, same era.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Cited alongside, same era.
Semantic Verification in Large Language Model-based Retrieval Augmented Generation. In Proceedings of the AAAI Symposium Series , Vol. 3. 188–192
Andreas Martin, Hans Friedrich Witschel, Maximilian Mandl, and Mona Stockhecke. 2024 · 2024
Cited alongside, same era.
QwQ: Reflect Deeply on the Boundaries of the Unknown
Qwen_Team. 2024 · 2024
Cited alongside, same era.
CBR-RAG: case-based reasoning for retrieval augmented generation in LLMs for legal question answering. In International Conference on Case-Based Reasoning . Springer, 445–460
Nirmalie Wiratunga, Ramitha Abeyratne, Lasal Jayawardena, Kyle Martin, Stewart Massie, Ikechukwu Nkisi-Orji, Ruvan Weerasinghe, Anne Liret, and Bruno Fleisch. 2024 · 2024
Cited alongside, same era.
When search engine services meet large language models: visions and challenges
Haoyi Xiong, Jiang Bian, Yuchen Li, Xuhong Li, Mengnan Du, Shuaiqiang Wang, Dawei Yin, and Sumi Helal. 2024 · 2024
Cited alongside, same era.
Crud-rag: A comprehensive chinese benchmark for retrieval-augmented generation of large language models
Yuanjie Lyu, Zhiyu Li, Simin Niu, Feiyu Xiong, Bo Tang, Wenjin Wang, Hao Wu, Huanyong Liu, Tong Xu, and Enhong Chen. 2025 · 2025
Closest in time.
https://kimi.moonshot.cn/
MoonShot_AI. 2025 · 2025
Closest in time.
https://www.perplexity.ai/
Perplexity. 2025 · 2025
Closest in time.
Qwen2.5-Max: Exploring the Intelligence of Large-scale MoE Model
Qwen_Team. 2025 · 2025
Closest in time.
https://yuanbao.tencent.com/chat/
Tencent. 2025 · 2025
Closest in time.
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. 2025 · 2025
Closest in time.
https://grok.com/
xAI. 2025 · 2025
Closest in time.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2025
Closest in time.