Fetching the paper…
Reading the bibliography…
Beyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios.
MLench: benchmarking Machine Learning Services Against Human Experts
Liu, Y.; Zhang, H.; Zeng, L.; Wu, W.; and Zhang, C. 2018 · 2018
Earlier work this paper cites.
Evaluating Large Language Models trained on Code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H. P. D. O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. 2021 · 2021
Earlier work this paper cites.
Measuring Coding Challenge Competence with Apps
Hendrycks, D.; Basart, S.; Kadavath, S.; Mazeika, M.; Arora, A.; Guo, E.; Burns, C.; Puranik, S.; He, H.; Song, D.; et al. 2021 · 2021
Earlier work this paper cites.
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
Lu, S.; Guo, D.; Ren, S.; Huang, J.; Svyatkovskiy, A.; Blanco, A.; Clement, C.; Drain, D.; Jiang, D.; Tang, D.; et al. 2021 · 2021
Earlier work this paper cites.
Jigsaw: Large Language Models Meet Program Synthesis
Jain, N.; Vaidyanath, S.; Iyer, A.; Natarajan, N.; Parthasarathy, S.; Rajamani, S.; and Sharma, R. 2022 · 2022
Earlier work this paper cites.
Competition-level Code Generation with Alphacode
Li, Y.; Choi, D.; Chung, J.; Kushman, N.; Schrittwieser, J.; Leblond, R.; Eccles, T.; Keeling, J.; Gimeno, F.; Dal Lago, A.; et al. 2022 · 2022
Earlier work this paper cites.
Execution-based Evaluation for Open-domain Code Generation
Wang, Z.; Zhou, S.; Fried, D.; and Neubig, G. 2022 · 2022
Earlier work this paper cites.
Crosscodeeval: A Diverse and Multilingual Benchmark for Cross-file Code Completion
Ding, Y.; Wang, Z.; Ahmad, W.; Ding, H.; Tan, M.; Jain, N.; Ramanathan, M. K.; Nallapati, R.; Bhatia, P.; Roth, D.; et al. 2023 · 2023
Earlier work this paper cites.
Classeval: A Manually-crafted Benchmark for Evaluating LLMs on Class-level Code Generation
Du, X.; Liu, M.; Wang, K.; Wang, H.; Liu, J.; Chen, Y.; Feng, J.; Sha, C.; Peng, X.; and Lou, Y. 2023 · 2023
Earlier work this paper cites.
MLAgentBench: Evaluating language Agents on Machine Learning Experimentation
Huang, Q.; Vora, J.; Liang, P.; and Leskovec, J. 2023 · 2023
Earlier work this paper cites.
Swe-bench: Can Language Models Resolve Real-World GitHub Issues??
Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K. 2023 · 2023
Earlier work this paper cites.
DS-1000: A natural and reliable benchmark for data science code generation
Lai, Y.; Li, C.; Wang, Y.; Zhang, T.; Zhong, R.; Zettlemoyer, L.; Yih, W.-t.; Fried, D.; Wang, S.; and Yu, T. 2023 · 2023
Earlier work this paper cites.
Api-bank: A Comprehensive Benchmark for Tool-augmented LLMs
Li, M.; Zhao, Y.; Yu, B.; Song, F.; Li, H.; Yu, H.; Li, Z.; Huang, F.; and Li, Y. 2023 · 2023
Earlier work this paper cites.
Repobench: Benchmarking Repository-level Code Auto-completion Systems
Liu, T.; Xu, C.; and McAuley, J. 2023 · 2023
Earlier work this paper cites.
Gitagent: Facilitating Autonomous Agent with Github by Tool Extension
Lyu, B.; Cong, X.; Yu, H.; Yang, P.; Qin, Y.; Ye, Y.; Lu, Y.; Zhang, Z.; Yan, Y.; Lin, Y.; et al. 2023 · 2023
Earlier work this paper cites.
Tang, X.; Liu, Y.; Cai, Z.; Shao, Y.; Lu, J.; Zhang, Y.; Deng, Z.; Hu, H.; An, K.; Huang, R.; et al. 2023 · 2023
Earlier work this paper cites.
Toolcoder: Teach Code Generation Models to Use Api Search Tools
Zhang, K.; Zhang, H.; Li, G.; Li, J.; Li, Z.; and Jin, Z. 2023 · 2023
Earlier work this paper cites.
Claude 3.5 Sonnet
Anthropic. 2024 · 2024
Cited alongside, same era.
Large language Models Empowered Agent-based Modeling and Simulation: A Survey and Perspectives
Gao, C.; Lan, X.; Li, N.; Yuan, Y.; Ding, J.; Zhou, Z.; Xu, F.; and Li, Y. 2024 · 2024
Cited alongside, same era.
Ishibashi, Y.; and Nishimura, Y. 2024 · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N.; Han, K.; Gu, A.; Li, W.-D.; Yan, F.; Zhang, T.; Wang, S.; Solar-Lezama, A.; Sen, K.; and Stoica, I. 2024 · 2024
Cited alongside, same era.
Devbench: A Comprehensive Benchmark for Software Development
Li, B.; Wu, W.; Tang, Z.; Shi, L.; Yang, J.; Li, J.; Yao, S.; Qian, C.; Hui, B.; Zhang, Q.; et al. 2024 · 2024
Cited alongside, same era.
xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
Chen, K.; Ren, Y.; Liu, Y.; Hu, X.; Tian, H.; Xie, T.; Liu, F.; Zhang, H.; Liu, H.; Gong, Y.; et al. 2025 · 2025
Closest in time.
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
Dong, P.; Tang, Z.; Liu, X.; Li, L.; Chu, X.; and Li, B. 2025 · 2025
Closest in time.
Fiverr Freelance Services
Fiverr. 2025 · 2025
Closest in time.
Freelancer Marketplace
Freelancer. 2025 · 2025
Closest in time.
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gemini Team, G. 2025 · 2025
Closest in time.
GitTaskBench: Anonymous GitHub Repository
GitTaskBench. 2025 · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. 2024 · 2024
Cited alongside, same era.
Ni, Z.; Li, Y.; Yang, N.; Shen, D.; Lv, P.; and Dong, D. 2024 · 2024
Cited alongside, same era.
Hello GPT-4o
OpenAI. 2024 · 2024
Cited alongside, same era.
Executable Code Actions Elicit Better LLM Agents
Wang, X.; Chen, Y.; Yuan, L.; Zhang, Y.; Li, Y.; Peng, H.; and Ji, H. 2024 · 2024
Cited alongside, same era.
Swe-agent: Agent-computer interfaces enable automated software engineering
Yang, J.; Jimenez, C.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; and Press, O. 2024 · 2024
Cited alongside, same era.
Ye, J.; Li, G.; Gao, S.; Huang, C.; Wu, Y.; Li, S.; Fan, X.; Dou, S.; Zhang, Q.; Gui, T.; et al. 2024 · 2024
Cited alongside, same era.
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Yu, Z.; Zhao, Y.; Cohan, A.; and Zhang, X.-P. 2024 · 2024
Cited alongside, same era.
Closest in time.
WeirdML Benchmark
Ihle, H. T. 2025 · 2025
Closest in time.
Artificial intelligence index report 2025
Maslej, N.; Fattorini, L.; Perrault, R.; Gil, Y.; Parli, V.; Kariuki, N.; Capstick, E.; Reuel, A.; Brynjolfsson, E.; Etchemendy, J.; et al. 2025 · 2025
Closest in time.
The Agentic Imperative Series Part 5: Manus and AutoGen
Masood, A. 2024 · 2025
Closest in time.
Llama 3.3
Meta AI. 2025 · 2025
Closest in time.
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
Miserendino, S.; Wang, M.; Patwardhan, T.; and Heidecke, J. 2025 · 2025
Closest in time.
PaperBench: Evaluating AI’s Ability to Replicate AI Research
Starace, G.; Jaffe, O.; Sherburn, D.; Aung, J.; Chan, J. S.; Maksin, L.; Dias, R.; Mays, E.; Kinsella, B.; Thompson, W.; et al. 2025 · 2025
Closest in time.
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
Tang, Z.; Liu, X.; Wang, Q.; Dong, P.; He, B.; Chu, X.; and Li, B. 2025 · 2025
Closest in time.
Upwork Online Freelance Marketplace
Upwork. 2025 · 2025
Closest in time.
Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; et al. 2025 · 2025
Closest in time.
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
Zheng, Z.; Cheng, Z.; Shen, Z.; Zhou, S.; Liu, K.; He, H.; Li, D.; Wei, S.; Hao, H.; Yao, J.; et al. 2025 · 2025
Closest in time.