Fetching the paper…
Reading the bibliography…
Recent reports claim that large language models (LLMs) now outperform elite humans in competitive programming.
Mathematics and games
Patrick M Grundy · 1939
Earlier work this paper cites.
Codeforces, 2010
Mike Mirzayanov · 2010
Earlier work this paper cites.
New features: friends, tags and more, 2011
Maxim Shipko · 2011
Earlier work this paper cites.
Interactive Problems: Guide for Participants, 2015
Codeforces · 2015
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Earlier work this paper cites.
Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion
Yangruibo Ding, Zijian Wang, Wasi Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, et al · 2023
Earlier work this paper cites.
Competition-level problems are effective llm evaluators
Yiming Huang, Zhenghao Lin, Xiao Liu, Yeyun Gong, Shuai Lu, Fangyu Lei, Yaobo Liang, Yelong Shen, Chen Lin, Nan Duan, et al · 2023
Earlier work this paper cites.
Model card addendum: Claude 3.5 sonnet
Anthropic · 2024
Earlier work this paper cites.
Claude 3.5 Sonnet, 2024
Anthropic · 2024
Earlier work this paper cites.
Mhpp: Exploring the capabilities and limitations of language models beyond basic code generation
Jianbo Dai, Jianqiao Lu, Yunlong Feng, Dong Huang, Guangtao Zeng, Rongju Ruan, Ming Cheng, Haochen Tan, and Zhijiang Guo · 2024
Cited alongside, same era.
Model card: Deepseek v3
DeepSeek AI · 2024
Cited alongside, same era.
Evaluating the performance of large language models in competitive programming: A multi-year, multi-grade analysis
Adrian Marius Dumitran, Adrian Cǎtǎlin Badea, and Stefan-Gabriel Muscalu · 2024
Cited alongside, same era.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, et al · 2024
Later among the works it cites.
URL https://icpc.global/worldfinals/fact-sheet/ICPC-Fact-Sheet.pdf
ICPC Fact Sheet, 2025 · 2025
Closest in time.
Official website of the china collegiate programming contest (ccpc), 2025
China Collegiate Programming Contest · 2025
Closest in time.
Model card: Qwen‑max (qwen 2.5 max)
Alibaba Cloud · 2025
Closest in time.
Model card: Gemini 2.0 flash reasoning
Google DeepMind · 2025
Closest in time.
Model card: Gemini 2.5 pro
Google DeepMind · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica · 2024
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan · 2024
Cited alongside, same era.
Model card: Meta llama 3.1 405b instruct
Meta AI · 2024
Cited alongside, same era.
System card: Gpt‑4o
OpenAI · 2024
Cited alongside, same era.
Can language models solve olympiad programming?
Quan Shi, Michael Tang, Karthik Narasimhan, and Shunyu Yao · 2024
Cited alongside, same era.
Learning task decomposition to assist humans in competitive programming
Jiaxin Wen, Ruiqi Zhong, Pei Ke, Zhihong Shao, Hongning Wang, and Minlie Huang · 2024
Cited alongside, same era.
Evaluating the smooth control of attribute intensity in text generation with llms
Shang Zhou, Feng Yao, Chengyu Dong, Zihan Wang, and Jingbo Shang · 2024
Cited alongside, same era.
URL https://github.com/openai/human-eval
HumanEval: Hand-Written Evaluation Set
Cited in the paper.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Llm-pros: Analyzing large language models’ performance in competitive problem solving
Md Sifat Hossain, Anika Tabassum, Md Fahim Arefin, and Tarannum Shaila Zaman · 2025
Closest in time.
Codeelo: Benchmarking competition-level code generation of llms with human-comparable elo ratings
Shanghaoran Quan, Jiaxi Yang, Bowen Yu, Bo Zheng, Dayiheng Liu, An Yang, Xuancheng Ren, Bofei Gao, Yibo Miao, Yunlong Feng, et al · 2025
Closest in time.
Model card: Llama 4 maverick 17b instruct
Unsloth (community release) · 2025
Closest in time.
Leetcodedataset: A temporal dataset for robust evaluation and efficient training of code llms
Yunhui Xia, Wei Shen, Yan Wang, Jason Klein Liu, Huifeng Sun, Siyue Wu, Jian Hu, and Xiaolong Xu · 2025
Closest in time.