Fetching the paper…
Reading the bibliography…
The increasing adoption of agentic workflows across diverse domains brings a critical need to scalably and systematically evaluate the complex traces these systems generate.
Back to the future: Towards explainable temporal reasoning with large language models
Chenhan Yuan, Qianqian Xie, Jimin Huang, and Sophia Ananiadou. 2024 · 1974
Earlier work this paper cites.
The limits of automatic summarisation according to ROUGE
Natalie Schluter. 2017 · 2017
Earlier work this paper cites.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Earlier work this paper cites.
What will it take to fix benchmarking in natural language understanding?
Samuel R Bowman and George E Dahl. 2021 · 2021
Earlier work this paper cites.
A fine-grained analysis of BERTScore
Michael Hanna and Ondřej Bojar. 2021 · 2021
Earlier work this paper cites.
Program synthesis with large language models
Augustus Odena, Charles Sutton, David Martin Dohan, Ellen Jiang, Henryk Michalewski, Jacob Austin, Maarten Paul Bosma, Maxwell Nye, Michael Terry, and Quoc V. Le. 2021 · 2021
Earlier work this paper cites.
Llm as os, agents as apps: Envisioning aios, agents and the aios-agent ecosystem
Yingqiang Ge, Yujie Ren, Wenyue Hua, Shuyuan Xu, Juntao Tan, and Yongfeng Zhang. 2023 · 2023
Earlier work this paper cites.
Jiayan Guo, Lun Du, Hengyu Liu, Mengyu Zhou, Xinyi He, and Shi Han. 2023 · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Earlier work this paper cites.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2023 · 2023
Earlier work this paper cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023 · 2023
Earlier work this paper cites.
Laser: Llm agent with state-space exploration for web navigation
Kaixin Ma, Hongming Zhang, Hongwei Wang, Xiaoman Pan, Wenhao Yu, and Dong Yu. 2023 · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants
Gregoire Mialon, Clementine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023 · 2023
Earlier work this paper cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al. 2023 · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023 · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Swe-bench+: Enhanced coding benchmark for llms
Reem Aleithan, Haoran Xue, Mohammad Mahdi Mohajer, Elijah Nnorom, Gias Uddin, and Song Wang. 2024 · 2024
Earlier work this paper cites.
Assessing and verifying task utility in LLM-powered applications
Negar Arabzadeh, Siqing Huo, Nikhil Mehta, Qingyun Wu, Chi Wang, Ahmed Hassan Awadallah, Charles L. A. Clarke, and Julia Kiseleva. 2024 · 2024
Earlier work this paper cites.
Mle-bench: Evaluating machine learning agents on machine learning engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal A. Patwardhan, Lil ian Weng, and Aleksander Mkadry. 2024 · 2024
Earlier work this paper cites.
Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark
Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. 2024 · 2024
Earlier work this paper cites.
Gamebench: Evaluating strategic reasoning abilities of llm agents
Anthony Costarelli, Mat Allen, Roman Hauksson, Grace Sodunke, Suhas Hariharan, Carlson Cheng, Wenjie Li, Joshua Clymer, and Arjun Yadav. 2024 · 2024
Earlier work this paper cites.
Contextualizing argument quality assessment with relevant knowledge
Darshan Deshpande, Zhivar Sourati, Filip Ilievski, and Fred Morstatter. 2024b · 2024
Earlier work this paper cites.
Evaluating large language models in class-level code generation
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou. 2024 · 2024
Earlier work this paper cites.
Imprompter: Tricking llm agents into improper tool use
Xiaohan Fu, Shuheng Li, Zihan Wang, Yihao Liu, Rajesh K Gupta, Taylor Berg-Kirkpatrick, and Earlence Fernandes. 2024 · 2024
Earlier work this paper cites.
Do llms estimate uncertainty well in instruction-following?
Juyeon Heo, Miao Xiong, Christina Heinze-Deml, and Jaya Narain. 2024 · 2024
Earlier work this paper cites.
Rag and rau: A survey on retrieval-augmented language model in natural language processing
Yucheng Hu and Yuxing Lu. 2024 · 2024
Earlier work this paper cites.
open deep research: An open-source replication of openai’s deep research agent
Hugging Face. 2024 · 2024
Earlier work this paper cites.
SWE-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024 · 2024
Earlier work this paper cites.
One thousand and one pairs: A" novel" challenge for long-context language models
Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2024 · 2024
Earlier work this paper cites.
Prometheus 2: An open source language model specialized in evaluating other language models
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024 · 2024
Earlier work this paper cites.
Spectool: A benchmark for characterizing errors in tool-use llms
Shirley Kokane, Ming Zhu, Tulika Awalgaonkar, Jianguo Zhang, Thai Hoang, Akshara Prabhakar, Zuxin Liu, Tian Lan, Liangwei Yang, Juntao Tan, et al. 2024 · 2024
Earlier work this paper cites.
Spotting AI‘s touch: Identifying LLM-paraphrased spans in text
Yafu Li, Zhilin Wang, Leyang Cui, Wei Bi, Shuming Shi, and Yue Zhang. 2024 · 2024
Earlier work this paper cites.
Coarse-to-fine highlighting: Reducing knowledge hallucination in large language models
Qitan Lv, Jie Wang, Hanzhu Chen, Bin Li, Yongdong Zhang, and Feng Wu. 2024 · 2024
Cited alongside, same era.
Llmparser: An exploratory study on using large language models for log parsing
Zeyang Ma, An Ran Chen, Dong Jae Kim, Tse-Husn Chen, and Shaowei Wang. 2024c · 2024
Cited alongside, same era.
Dynasaur: Large language agents beyond predefined actions
Dang Nguyen, Viet Dac Lai, Seunghyun Yoon, Ryan A. Rossi, Handong Zhao, Ruiyi Zhang, Puneet Mathur, Nedim Lipka, Yu Wang, Trung Bui, Franck Dernoncourt, and Tianyi Zhou. 2024 · 2024
Cited alongside, same era.
Openmanus-rl: An open-source rl environment for evaluating multimodal llms on scientific reasoning
OpenManus. 2024 · 2024
Cited alongside, same era.
Training software engineering agents and verifiers with swe-gym
Jiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, and Yizhe Zhang. 2024 · 2024
Judgelrm: Large reasoning models as a judge
Nuo Chen, Zhiyuan Hu, Qingyun Zou, Jiaying Wu, Qian Wang, Bryan Hooi, and Bingsheng He. 2025 · 2025
Closest in time.
Gemini model thinking updates: March 2025
Google DeepMind. 2025 · 2025
Closest in time.
Interactive debugging and steering of multi-agent ai systems
Will Epperson, Gagan Bansal, Victor Dibia, Adam Fourney, Jack Gerrits, Erkang Zhu, and Saleema Amershi. 2025 · 2025
Closest in time.
Fiction.livebench (april 6, 2025)
Ficlive. 2025 · 2025
Closest in time.
Synergizing rag and reasoning: A systematic review
Yunfan Gao, Yun Xiong, Yijie Zhong, Yuxi Bi, Ming Xue, and Haofen Wang. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gorilla: Large language model connected with massive APIs
Shishir G Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. 2024 · 2024
Cited alongside, same era.
ChatDev: Communicative agents for software development
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Cited alongside, same era.
Exploring llm-based agents for root cause analysis
Devjeet Roy, Xuchao Zhang, Rashi Bhave, Chetan Bansal, Pedro Las-Casas, Rodrigo Fonseca, and Saravan Rajmohan. 2024b · 2024
Cited alongside, same era.
Zhuocheng Shen. 2024 · 2024
Cited alongside, same era.
Structuredrag: Json response formatting with large language models
Connor Shorten, Charles Pierse, Thomas Benjamin Smith, Erika Cardenas, Akanksha Sharma, John Trengrove, and Bob van Luijt. 2024 · 2024
Cited alongside, same era.
Chain of thoughtlessness? an analysis of cot in planning
Kaya Stechly, Karthik Valmeekam, and Subbarao Kambhampati. 2024 · 2024
Cited alongside, same era.
Table meets llm: Can large language models understand structured table data? a benchmark and empirical study
Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. 2024 · 2024
Cited alongside, same era.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025 · 2025
Closest in time.
L4: Diagnosing large-scale llm training failures via automated log analysis
Zhihan Jiang, Junjie Huang, Zhuangbin Chen, Yichen Li, Guangba Yu, Cong Feng, Yongqiang Yang, Zengyin Yang, and Michael R. Lyu. 2025 · 2025
Closest in time.
A survey of frontiers in llm reasoning: Inference scaling, learning to reason, and agentic systems
Zixuan Ke, Fangkai Jiao, Yifei Ming, Xuan-Phi Nguyen, Austin Xu, Do Xuan Long, Minzhi Li, Chengwei Qin, Peifeng Wang, Silvio Savarese, et al. 2025 · 2025
Closest in time.
The BiGGen bench: A principled benchmark for fine-grained evaluation of language models with language models
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2025 · 2025
Closest in time.
Acpbench: Reasoning about action, change, and planning
Harsha Kokel, Michael Katz, Kavitha Srinivas, and Shirin Sohrabi. 2025 · 2025
Closest in time.
Llms get lost in multi-turn conversation
Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. 2025 · 2025
Closest in time.
Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists
Yukyung Lee, Joonghoon Kim, Jaehee Kim, Hyowon Cho, Jaewook Kang, Pilsung Kang, and Najoung Kim. 2025 · 2025
Closest in time.
Llama 4: Advancing multimodal intelligence
Meta AI. 2025 · 2025
Closest in time.
Toolfuzz–automated agent tool testing
Ivan Milev, Mislav Balunović, Maximilian Baader, and Martin Vechev. 2025 · 2025
Closest in time.
Beyond black-box benchmarking: Observability, analytics, and optimization of agentic systems
Dany Moshkovich, Hadar Mulian, Sergey Zeltyn, Natti Eder, Inna Skarbovsky, and Roy Abitbol. 2025 · 2025
Closest in time.
Governance in agentic workflows: Leveraging llms as oversight agents
Imran Nasim. 2025 · 2025
Closest in time.
Introducing deep research
OpenAI. 2024 · 2025
Closest in time.
Introducing GPT-4.1
OpenAI. 2025a · 2025
Closest in time.
Introducing O1: A state-of-the-art multimodal ai model
OpenAI. 2025b · 2025
Closest in time.
Introducing o3 and o4-mini
OpenAI. 2025c · 2025
Closest in time.
Introducing o3-mini: A smaller, faster and more cost-effective model
OpenAI. 2025d · 2025
Closest in time.
OpenTelemetry — opentelemetry.io
OpenTelemetry. 2025 · 2025
Closest in time.
Modeling statistical risk in ai products
Patronus AI. 2025 · 2025
Closest in time.
Long Phan et al. 2025 · 2025
Closest in time.
Learning to plan & reason for evaluation with thinking-llm-as-a-judge
Swarnadeep Saha, Xian Li, Marjan Ghazvininejad, Jason Weston, and Tianlu Wang. 2025 · 2025
Closest in time.
Bright: A realistic and challenging benchmark for reasoning-intensive retrieval
Hongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi, Niklas Muennighoff, Han yu Wang, Haisu Liu, Quan Shi, Zachary S. Siegel, Michael Tang, Ruoxi Sun, Jinsung Yoon, Sercan O. Arik, Danqi Chen, and Tao Yu. 2025 · 2025
Closest in time.
Does context matter? contextualjudgebench for evaluating llm-based judges in contextual settings
Austin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz, and Shafiq Joty. 2025 · 2025
Closest in time.
Survey on evaluation of llm-based agents
Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. 2025 · 2025
Closest in time.
Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space
Yong Zhao, Kai Xu, Zhengqiu Zhu, Yue Hu, Zhiheng Zheng, Yingfeng Chen, Yatai Ji, Chen Gao, Yong Li, and Jincai Huang. 2025 · 2025
Closest in time.
Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, and Shafiq Joty. 2025 · 2025
Closest in time.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2025 · 2025
Closest in time.