Fetching the paper…
Reading the bibliography…
Failure attribution in LLM multi-agent systems-identifying the agent and step responsible for task failures-provides crucial clues for systems debugging but remains underexplored and labor-intensive.
Combining forecasts: A review and annotated bibliography
Clemen, R. T · 1989
Earlier work this paper cites.
Evaluation thesaurus
Scriven, M · 1991
Earlier work this paper cites.
Hyper-parameter optimization: A review of algorithms and applications
Yu, T. and Zhu, H · 2020
Earlier work this paper cites.
Towards reasoning in large language models: A survey
Huang, J. and Chang, K. C.-C · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z · 2023
Earlier work this paper cites.
Gptscore: Evaluate as you desire
Fu, J., Ng, S.-K., Jiang, Z., and Liu, P · 2023
Earlier work this paper cites.
Metagpt: Meta programming for multi-agent collaborative framework
Hong, S., Zheng, X., Chen, J., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., et al · 2023
Earlier work this paper cites.
Large language models cannot self-correct reasoning yet
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Earlier work this paper cites.
Let’s verify step by step
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C · 2023
Cited alongside, same era.
Gaia: a benchmark for general ai assistants
Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., and Scialom, T · 2023
Cited alongside, same era.
Selfcheck: Using llms to zero-shot check their own step-by-step reasoning
Miao, N., Teh, Y. W., and Rainforth, T · 2023
Cited alongside, same era.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Wang, P., Li, L., Shao, Z., Xu, R., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., and Wang, C · 2023
Cited alongside, same era.
Adaptive in-conversation team building for language model agents
Song, L., Liu, J., Zhang, J., Zhang, S., Luo, A., Wang, S., Wu, Q., and Wang, C · 2024
Later among the works it cites.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., and Hupkes, D · 2024
Later among the works it cites.
A field guide to automatic evaluation of llm-generated summaries
van Schaik, T. A. and Pugh, B · 2024
Later among the works it cites.
Wang, F., Zhang, Z., Zhang, X., Wu, Z., Mo, T., Lu, Q., Wang, W., Li, R., Xu, J., Tang, X., et al · 2024
Later among the works it cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Cited alongside, same era.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2024
Cited alongside, same era.
Magentic-one: A generalist multi-agent system for solving complex tasks
Fourney, A., Bansal, G., Mozannar, H., Tan, C., Salinas, E., Niedtner, F., Proebsting, G., Bassman, G., Gerrits, J., Alber, J., et al · 2024
Cited alongside, same era.
Sciagents: Automating scientific discovery through multi-agent intelligent graph reasoning
Ghafarollahi, A. and Buehler, M. J · 2024
Cited alongside, same era.
Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., et al · 2024
Cited alongside, same era.
Language model preference evaluation with multiple weak evaluators
Hu, Z., Zhang, J., Xiong, Z., Ratner, A., Xiong, H., and Krishna, R · 2024
Cited alongside, same era.
Needle in the haystack for memory based large language models
Nelson, E., Kollias, G., Das, P., Chaudhury, S., and Dan, S · 2024
Cited alongside, same era.
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., et al · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Assistantbench: Can web agents solve realistic and time-consuming tasks?
Yoran, O., Amouyal, S. J., Malaviya, C., Bogin, B., Press, O., and Berant, J · 2024
Later among the works it cites.
Processbench: Identifying process errors in mathematical reasoning
Zheng, C., Zhang, Z., Zhang, B., Lin, R., Lu, K., Yu, B., Liu, D., Zhou, J., and Lin, J · 2024
Later among the works it cites.
Agent-as-a-judge: Evaluate agents with agents
Zhuge, M., Zhao, C., Ashley, D., Wang, W., Khizbullin, D., Xiong, Y., Liu, Z., Chang, E., Krishnamoorthi, R., Tian, Y., et al · 2024
Later among the works it cites.
Process reinforcement through implicit rewards
Cui, G., Yuan, L., Wang, Z., Wang, H., Li, W., He, B., Fan, Y., Yu, T., Xu, Q., Chen, W., et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning deepseek-ai
DeepSeek-AI · 2025
Closest in time.
Nemotron-research-tool-n1: Tool-using language models with reinforced reasoning
Zhang, S., Dong, Y., Zhang, J., Kautz, J., Catanzaro, B., Tao, A., Wu, Q., Yu, Z., and Liu, G · 2025
Closest in time.
A comprehensive survey of reward models: Taxonomy, applications, challenges, and future
Zhong, J., Shen, W., Li, Y., Gao, S., Lu, H., Chen, Y., Zhang, Y., Zhou, W., Gu, J., and Zou, L · 2025
Closest in time.