Fetching the paper…
Reading the bibliography…
Contemporary evaluation techniques are inadequate for agentic systems.
Binary codes capable of correcting deletions, insertions, and reversals
V Levenshtein · 1966
Earlier work this paper cites.
Monetary relationships: a view from Threadneedle Street
Charles Goodhart · 1976
Earlier work this paper cites.
Thirteen theorems in search of the truth
Bernard Grofman, Guillermo Owen, and Scott L Feld · 1983
Earlier work this paper cites.
Combining forecasts: A review and annotated bibliography
Robert T Clemen · 1989
Earlier work this paper cites.
Support-vector networks
Corinna Cortes · 1995
Earlier work this paper cites.
From data mining to knowledge discovery in databases
Usama Fayyad, Gregory Piatetsky-Shapiro, and Padhraic Smyth · 1996
Earlier work this paper cites.
Long short-term memory
S Hochreiter · 1997
Earlier work this paper cites.
Intelligent agents
Michael Wooldridge · 1999
Earlier work this paper cites.
Crisp-dm: Towards a standard process model for data mining
Rüdiger Wirth and Jochen Hipp · 2000
Earlier work this paper cites.
The robust beauty of majority rules in group decisions
Reid Hastie and Tatsuya Kameda · 2005
Earlier work this paper cites.
Intuitions about combining opinions: Misappreciation of the averaging principle
Richard P Larrick and Jack B Soll · 2006
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Trevor Hastie, Robert Tibshirani, and Jerome H. Friedman · 2009
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Constructive multiple-choice testing system
Jooyong Park · 2010
Earlier work this paper cites.
Learning to generalize from sparse and underspecified rewards
Rishabh Agarwal, Chen Liang, Dale Schuurmans, and Mohammad Norouzi · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N Reimers · 2019
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Automl: A survey of the state-of-the-art
Xin He, Kaiyong Zhao, and Xiaowen Chu · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al · 2021
Earlier work this paper cites.
LangChain
Harrison Chase · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Earlier work this paper cites.
Multipl-e: a scalable and polyglot approach to benchmarking neural code generation
Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, et al · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Earlier work this paper cites.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Earlier work this paper cites.
Auto-gpt
Significant Gravitas · 2023
Earlier work this paper cites.
Fixeval: Execution-based evaluation of program fixes for competitive programming problems
Md Mahim Anjum Haque · 2023
Cited alongside, same era.
Agentcoder: Multi-agent-based code generation with iterative testing and optimisation
Dong Huang, Qingwen Bu, Jie M Zhang, Michael Luck, and Heming Cui · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
Dspy: Compiling declarative language model calls into self-improving pipelines
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, et al · 2023
Cited alongside, same era.
Ds-1000: A natural and reliable benchmark for data science code generation
From llms to llm-based agents for software engineering: A survey of current, challenges and future
Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan, Bo Li, and Huaming Chen · 2024
Closest in time.
Dsbench: How far are data science agents to becoming data science experts?
Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu · 2024
Closest in time.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried · 2024
Closest in time.
LangGraph
langchain ai · 2024
Closest in time.
Mlr-copilot: Autonomous machine learning research based on large language models agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-tau Yih, Daniel Fried, Sida Wang, and Tao Yu · 2023
Cited alongside, same era.
Camel: Communicative agents for” mind” exploration of large scale language model society
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Cited alongside, same era.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
Gpt-pilot: Your ai copilot for software development
Pythagora.io · 2023
Cited alongside, same era.
Taskweaver: A code-first agent framework
Bo Qiao, Liqun Li, Xu Zhang, Shilin He, Yu Kang, Chaoyun Zhang, Fangkai Yang, Hang Dong, Jue Zhang, Lu Wang, et al · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang · 2023
Cited alongside, same era.
Appagent: Multimodal agents as smartphone users
Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu · 2023
Cited alongside, same era.
Repocoder: Repository-level code completion through iterative retrieval and generation
Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen · 2023
Cited alongside, same era.
Ruochen Li, Teerth Patel, Qingyun Wang, and Xinya Du · 2024
Closest in time.
Large language model-based agents for software engineering: A survey
Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou · 2024
Closest in time.
Arena learning: Build data flywheel for llms post-training via simulated chatbot arena
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Qingwei Lin, Jianguang Lou, Shifeng Chen, Yansong Tang, and Weizhu Chen · 2024
Closest in time.
Code agents are state of the art software testers
Niels Mündler, Mark Niklas Müller, Jingxuan He, and Martin Vechev · 2024
Closest in time.
Autonomous evaluation and refinement of digital agents
Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr · 2024
Closest in time.
Hyperagent: Generalist software engineering agents to solve coding tasks at scale
Huy Nhat Phan, Phong X Nguyen, and Nghi DQ Bui · 2024
Closest in time.
Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment
Vyas Raina, Adian Liusie, and Mark Gales · 2024
Closest in time.
Lin Shi, Weicheng Ma, and Soroush Vosoughi · 2024
Closest in time.
Adaptive in-conversation team building for language model agents
Linxin Song, Jiale Liu, Jieyu Zhang, Shaokun Zhang, Ao Luo, Shijian Wang, Qingyun Wu, and Chi Wang · 2024
Closest in time.
Towards general computer control: A multimodal agent for red dead redemption ii as a case study
Weihao Tan, Ziluo Ding, Wentao Zhang, Boyu Li, Bohan Zhou, Junpeng Yue, Haochong Xia, Jiechuan Jiang, Longtao Zheng, Xinrun Xu, et al · 2024
Closest in time.
Magis: Llm-based multi-agent framework for github issue resolution
Wei Tao, Yucheng Zhou, Wenqiang Zhang, and Yu Cheng · 2024
Closest in time.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes · 2024
Closest in time.
Debugbench: Evaluating debugging capability of large language models
Runchu Tian, Yining Ye, Yujia Qin, Xin Cong, Yankai Lin, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Autodev: Automated ai-driven development
Michele Tufano, Anisha Agarwal, Jinu Jang, Roshanak Zilouchian Moghaddam, and Neel Sundaresan · 2024
Closest in time.
Agentless: Demystifying llm-based software engineering agents
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang · 2024
Closest in time.
Can large language model agents simulate human trust behaviors?
Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Kai Shu, Adel Bibi, Ziniu Hu, Philip Torr, Bernard Ghanem, and Guohao Li · 2024
Closest in time.
Llava-critic: Learning to evaluate multimodal models
Tianyi Xiong, Xiyao Wang, Dong Guo, Qinghao Ye, Haoqi Fan, Quanquan Gu, Heng Huang, and Chunyuan Li · 2024
Closest in time.
Crab: Cross-environment agent benchmark for multimodal language model agents
Tianqi Xu, Linyao Chen, Dai-Jie Wu, Yanjun Chen, Zecheng Zhang, Xiang Yao, Zhiqiang Xie, Yongchao Chen, Shilong Liu, Bochen Qian, et al · 2024
Closest in time.
Auto arena of llms: Automating llm evaluations with agent peer-battles and committee discussions
Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Deli Zhao, and Lidong Bing · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.
Symbolic learning enables self-evolving agents
Wangchunshu Zhou, Yixin Ou, Shengwei Ding, Long Li, Jialong Wu, Tiannan Wang, Jiamin Chen, Shuai Wang, Xiaohua Xu, Ningyu Zhang, et al · 2024
Closest in time.
Language agents as optimizable graphs
Mingchen Zhuge, Wenyi Wang, Louis Kirsch, Francesco Faccio, Dmitrii Khizbullin, and Jurgen Schmidhuber · 2024
Closest in time.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, et al · 2024
Closest in time.