Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) excel in code-related tasks like code generation, but benchmark evaluations often overlook task characteristics, such as difficulty.
A new measure of rank correlation
Maurice G Kendall. 1938 · 1938
Earlier work this paper cites.
A simple sequentially rejective multiple test procedure
Sture Holm. 1979 · 1979
Earlier work this paper cites.
K-sample Anderson–Darling tests
Fritz W Scholz and Michael A Stephens. 1987 · 1987
Earlier work this paper cites.
Item clusters and computerized adaptive testing: A case for testlets
Howard Wainer and Gerard L Kiely. 1987 · 1987
Earlier work this paper cites.
Individual comparisons by ranking methods
Frank Wilcoxon. 1992 · 1992
Earlier work this paper cites.
Table for conversion of Kendall’s Tau to Spearman’s Rho within the context of measures of magnitude of effect for meta-analysis
Andrew R Gilpin. 1993 · 1993
Earlier work this paper cites.
Interpretation of Kappa and B statistics measures of agreement
Sergio R Munoz and Shrikant I Bangdiwala. 1997 · 1997
Earlier work this paper cites.
The basics of item response theory
Frank B Baker. 2001 · 2001
Earlier work this paper cites.
Planning poker or how to avoid analysis paralysis while release planning
James Grenning. 2002 · 2002
Earlier work this paper cites.
The origins of testlet response theory – three alternatives
Howard Wainer, Eric T. Bradlow, and Xiaohui Wang. 2007 · 2007
Earlier work this paper cites.
Effect size estimates: current use, calculations, and interpretation
Catherine O Fritz, Peter E Morris, and Jennifer J Richler. 2012 · 2012
Earlier work this paper cites.
Building an evaluation scale using item response theory. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing , Vol. 2016. NIH Public Access, 648
John P Lalor, Hao Wu, and Hong Yu. 2016 · 2016
Earlier work this paper cites.
An application of item response theory to psychological test development
Cristian Zanon, Claudio S Hutz, Hanwook Henry Yoo, and Ronald K Hambleton. 2016 · 2016
Earlier work this paper cites.
hdbscan: Hierarchical density based clustering
Leland McInnes, John Healy, Steve Astels, et al · 2017
Earlier work this paper cites.
Metamorphic testing: A review of challenges and opportunities
Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu, Pak-Lok Poon, Dave Towey, TH Tse, and Zhi Quan Zhou. 2018 · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville. 2018 · 2018
Earlier work this paper cites.
Correlation coefficients: appropriate use and interpretation
Patrick Schober, Christa Boer, and Lothar A Schwarte. 2018 · 2018
Earlier work this paper cites.
β 3 \beta^{3} -IRT: A New Item Response Model and its Applications. In The 22nd International Conference on Artificial Intelligence and Statistics . PMLR, 1013–1021
Yu Chen, Telmo Silva Filho, Ricardo B Prudencio, Tom Diethe, and Peter Flach. 2019 · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2020
Earlier work this paper cites.
Item Response Theory for Efficient Human Evaluation of Chatbots. In Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems , Steffen Eger, Yang Gao, Maxime Peyrard, Wei Zhao, and Eduard Hovy (Eds.). Association for Computational Linguistics, Online, 21–33
João Sedoc and Lyle Ungar. 2020 · 2020
Earlier work this paper cites.
Playing planning poker in crowds: human computation of software effort estimates. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 1–12
Mohammed Alhamed and Tim Storer. 2021 · 2021
Earlier work this paper cites.
Program Synthesis with Large Language Models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
A fine-grained analysis of BERTScore. In Proceedings of the Sixth Conference on Machine Translation . 507–517
Michael Hanna and Ondřej Bojar. 2021 · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al · 2021
Earlier work this paper cites.
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, et al · 2021
Earlier work this paper cites.
Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 4486–4503
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. 2021b · 2021
Earlier work this paper cites.
Comparing test sets with item response theory
Clara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, and Samuel R Bowman. 2021a · 2021
Earlier work this paper cites.
Comparing test sets with item response theory
Clara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, and Samuel R Bowman. 2021b · 2021
Earlier work this paper cites.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Cited alongside, same era.
Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li, Yuchen Tian, Ming Tan, Wasi Uddin Ahmad, Shiqi Wang, Qing Sun, Mingyue Shang, Sujan Kumar Gonugondla, Hantian Ding, Varun Kumar, Nathan Fulton, Arash Farahani, Siddhartha Jain, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Sudipta Sengupta, Dan Roth, and Bing Xiang. 2022 · 2022
Cited alongside, same era.
Predicting Difficulty and Discrimination of Natural Language Questions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 119–130
Matthew Byrd and Shashank Srivastava. 2022 · 2022
Cited alongside, same era.
BERTopic: Neural topic modeling with a class-based TF-IDF procedure
ClassEval Leaderboard
2024 · 2024
Closest in time.
EvalPlus Leaderboard
2024 · 2024
Closest in time.
Models on Hugging Face
2024 · 2024
Closest in time.
Rep-Package
2024 · 2024
Closest in time.
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, Benjamin L. Edelman, Zhaowei Zhang, Mario Günther, Anton Korinek, Jose Hernandez-Orallo, Lewis Hammond, Eric Bigelow, Alexander Pan, Lauro Langosco, Tomasz Korbak, Heidi Zhang, Ruiqi Zhong, Seán Ó hÉigeartaigh, Gabriel Recchia, Giulio Corsi, Alan Chan, Markus Anderljung, Lilian Edwards, Aleksandar Petrov, Christian Schroeder de Witt, Sumeet Ramesh Motwan, Yoshua Bengio, Danqi Chen, Philip H. S. Torr, Samuel Albanie, Tegan Maharaj, Jakob Foerster, Florian Tramer, He He, Atoosa Kasirzadeh, Yejin Choi, and David Krueger. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maarten Grootendorst. 2022 · 2022
Cited alongside, same era.
Competition-level code generation with AlphaCode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals. 2022 · 2022
Cited alongside, same era.
An empirical evaluation of GitHub copilot’s code suggestions. In Proceedings of the 19th International Conference on Mining Software Repositories . 1–5
Nhan Nguyen and Sarah Nadi. 2022 · 2022
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022a · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Rephrase and respond: Let large language models ask better questions for themselves
Yihe Deng, Weitong Zhang, Zixiang Chen, and Quanquan Gu. 2023 · 2023
Cited alongside, same era.
Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou. 2023 · 2023
Cited alongside, same era.
Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. 2023 · 2023
Cited alongside, same era.
A Performance Study of LLM-Generated Code on Leetcode. In Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering (Salerno, Italy) (EASE ’24) . Association for Computing Machinery, New York, NY, USA, 79–89
Tristan Coignion, Clément Quinton, and Romain Rouvoy. 2024 · 2024
Closest in time.
DeepSeek-AI. 2024 · 2024
Closest in time.
DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y Wu, YK Li, et al · 2024
Closest in time.
Hugging Face — huggingface.co
HuggingFace [n. d.] · 2024
Closest in time.
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024 · 2024
Closest in time.
A Survey on Large Language Models for Code Generation
Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. 2024 · 2024
Closest in time.
SWE-bench: Can Language Models Resolve Real-world Github Issues?. In The Twelfth International Conference on Learning Representations
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024 · 2024
Closest in time.
Using AI Assistants in Software Development: A Qualitative Study on Security Practices and Concerns
Jan H Klemmer, Stefan Albert Horstmann, Nikhil Patnaik, Cordelia Ludden, Cordell Burton Jr, Carson Powers, Fabio Massacci, Akond Rahman, Daniel Votipka, Heather Richter Lipford, et al · 2024
Closest in time.
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 13785–13816
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty, and Jimmy Huang. 2024 · 2024
Closest in time.
Exploring and Evaluating Hallucinations in LLM-Powered Code Generation
Fang Liu, Yang Liu, Lin Shi, Houkun Huang, Ruifeng Wang, Zhen Yang, and Li Zhang. 2024a · 2024
Closest in time.
StarCoder 2 and The Stack v2: The Next Generation
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, et al · 2024
Closest in time.
Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Timothy R. McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Paul Watters, and Malka N. Halgamuge. 2024 · 2024
Closest in time.
State of What Art? A Call for Multi-Prompt LLM Evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. 2024 · 2024
Closest in time.
An Empirical Study of the Non-determinism of ChatGPT in Code Generation
Shuyin Ouyang, Jie M. Zhang, Mark Harman, and Meng Wang. 2024 · 2024
Closest in time.
Systematic literature review of prompt engineering patterns in software engineering. In 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC) . IEEE, 670–675
Yuya Sasaki, Hironori Washizaki, Jialong Li, Dominik Sander, Nobukazu Yoshioka, and Yoshiaki Fukazawa. 2024 · 2024
Closest in time.
Quality Assessment of Prompts Used in Code Generation
Mohammed Latif Siddiq, Simantika Dristi, Joy Saha, and Joanna C. S. Santos. 2024 · 2024
Closest in time.
Bugs in large language models generated code
Florian Tambon, Arghavan Moradi Dakhel, Amin Nikanjam, Foutse Khomh, Michel C Desmarais, and Giuliano Antoniol. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
LLMs in Web-Development: Evaluating LLM-Generated PHP code unveiling vulnerabilities and limitations
Rebeka Tóth, Tamas Bisztray, and László Erdodi. 2024 · 2024
Closest in time.
Validating LLM-Generated Programs with Metamorphic Prompt Testing
Xiaoyin Wang and Dakai Zhu. 2024 · 2024
Closest in time.
Top Leaderboard Ranking= Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
Chunqiu Steven Xia, Yinlin Deng, and Lingming Zhang. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan. 2024 · 2024
Closest in time.
From Static Benchmarks to Adaptive Testing: Psychometrics in AI Evaluation
Yan Zhuang, Qi Liu, Yuting Ning, Weizhe Huang, Zachary A. Pardos, Patrick C. Kyllonen, Jiyun Zu, Qingyang Mao, Rui Lv, Zhenya Huang, Guanhao Zhao, Zheng Zhang, Shijin Wang, and Enhong Chen. 2024 · 2024
Closest in time.
A Study on the Pythonic Functional Constructs’ Understandability. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
Cyrine Zid, Fiorella Zampetti, Giuliano Antoniol, and Massimiliano Di Penta. 2024 · 2024
Closest in time.
Yi: Open Foundation Models by 01.AI
01. AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Guoyin Wang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yanpeng Li, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai. 2025 · 2025
Closest in time.