Fetching the paper…
Reading the bibliography…
Long-horizon reasoning in LLM-based agents often fails not from generative weakness but from insufficient verification of intermediate reasoning.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Earlier work this paper cites.
Self-verification improves few-shot clinical information extraction
Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley, Jianfeng Gao, and Hoifung Poon · 2023
Earlier work this paper cites.
SELFCHECKGPT: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark JF Gales · 2023
Earlier work this paper cites.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Earlier work this paper cites.
ReAct: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Earlier work this paper cites.
Adversarial multi-agent evaluation of large language models through iterative debates
Chaithanya Bandi and Abir Harrasse · 2024
Earlier work this paper cites.
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al · 2024
Earlier work this paper cites.
CHECKEMBED: Effective verification of LLM solutions to open-ended tasks
Maciej Besta, Lorenzo Paleari, Marcin Copik, Robert Gerstenberger, Ales Kubicek, Piotr Nyczyk, Patrick Iff, Eric Schreiber, Tanja Srindran, Tomasz Lehmann, et al · 2024
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch · 2024
Earlier work this paper cites.
From local to global: A graph RAG approach to query-focused summarization
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson · 2024
Earlier work this paper cites.
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal · 2024
Earlier work this paper cites.
Large language model based multi-agents: A survey of progress and challenges
Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang · 2024
Earlier work this paper cites.
Chinese simpleQA: A chinese factuality evaluation for large language models
Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, et al · 2024
Cited alongside, same era.
Diversity of thought elicits stronger reasoning capabilities in multi-agent debate frameworks
Mahmood Hegazy · 2024
Cited alongside, same era.
LongRAG: Enhancing retrieval-augmented generation with long-context LLMs
Ziyan Jiang, Xueguang Ma, and Wenhu Chen · 2024
Cited alongside, same era.
When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs
Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang · 2024
Cited alongside, same era.
Internal consistency and self-feedback in large language models: A survey
From LLM reasoning to autonomous AI agents: A comprehensive review
Mohamed Amine Ferrag, Norbert Tihanyi, and Merouane Debbah · 2025
Closest in time.
Plan verification for LLM-based embodied task completion agents
Ananth Hariharan, Vardhan Dongre, Dilek Hakkani-Tür, and Gokhan Tur · 2025
Closest in time.
Deep research agents: A systematic examination and roadmap
Yuxuan Huang, Yihang Chen, Haozheng Zhang, Kang Li, Huichi Zhou, Meng Fang, Linyi Yang, Xiaoguang Li, Lifeng Shang, Songcen Xu, et al · 2025
Closest in time.
ReflAct: World-grounded decision making in LLM agents via goal-state reflection
Jeonghye Kim, Sojeong Rhee, Minbeom Kim, Dohyung Kim, Sangmook Lee, Youngchul Sung, and Kyomin Jung · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xun Liang, Shichao Song, Zifan Zheng, Hanyu Wang, Qingchen Yu, Xunkai Li, Rong-Hua Li, Yi Wang, Zhonghao Wang, Feiyu Xiong, et al · 2024
Cited alongside, same era.
GAIA: a benchmark for general AI assistants
Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2024
Cited alongside, same era.
Reasoning with large language models, a survey
Aske Plaat, Annie Wong, Suzan Verberne, Joost Broekens, Niki van Stein, and Thomas Bäck · 2024
Cited alongside, same era.
MiniCheck: Efficient fact-checking of LLMs on grounding documents
Liyan Tang, Philippe Laban, and Greg Durrett · 2024
Cited alongside, same era.
On the brittle foundations of ReAct prompting for agentic large language models
Mudit Verma, Siddhant Bhambri, and Subbarao Kambhampati · 2024
Cited alongside, same era.
Gemini deep research: Automated research assistant
Google AI · 2025
Cited alongside, same era.
Kimi-researcher: End-to-end RL training for emerging agentic capabilities
Moonshot AI · 2025
Cited alongside, same era.
Why do multi-agent LLM systems fail?
Mert Cemri, Melissa Z Pan, Shuyi Yang, Lakshya A Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, et al · 2025
Cited alongside, same era.
Kuan Li, Zhongwang Zhang, Huifeng Yin, Liwen Zhang, Litu Ou, Jialong Wu, Wenbiao Yin, Baixuan Li, Zhengwei Tao, Xinyu Wang, et al · 2025
Closest in time.
Large language model agent: A survey on methodology, applications and challenges
Junyu Luo, Weizhi Zhang, Ye Yuan, Yusheng Zhao, Junwei Yang, Yiyang Gu, Bohan Wu, Binqi Chen, Ziyue Qiao, Qingqing Long, et al · 2025
Closest in time.
SelfCheckAgent: Zero-resource hallucination detection in generative large language models
Diyana Muhammed, Gollam Rabby, and Sören Auer · 2025
Closest in time.
Mihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, et al · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al · 2025
Closest in time.
WebResearcher: Unleashing unbounded reasoning capability in long-horizon agents
Zile Qiao, Guoxin Chen, Xuanzhong Chen, Donglei Yu, Wenbiao Yin, Xinyu Wang, Zhen Zhang, Baixuan Li, Huifeng Yin, Kuan Li, et al · 2025
Closest in time.
AWorld: Orchestrating the training recipe for agentic AI
Chengyue Yu, Siyuan Lu, Chenyi Zhuang, Dong Wang, Qintong Wu, Zongyue Li, Runsheng Gan, Chunfeng Wang, Siqi Hou, Gaochi Huang, et al · 2025
Closest in time.
Dyna-Think: Synergizing reasoning, acting, and world model simulation in ai agents
Xiao Yu, Baolin Peng, Ruize Xu, Michel Galley, Hao Cheng, Suman Nath, Jianfeng Gao, and Zhou Yu · 2025
Closest in time.
AgentOrchestra: A hierarchical multi-agent framework for general-purpose task solving
Wentao Zhang, Liang Zeng, Yuzhen Xiao, Yongcong Li, Ce Cui, Yilei Zhao, Rui Hu, Yang Liu, Yahui Zhou, and Bo An · 2025
Closest in time.
A survey on the memory mechanism of large language model-based agents
Zeyu Zhang, Quanyu Dai, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen · 2025
Closest in time.