Fetching the paper…
Reading the bibliography…
Since the emergence of Large Language Models (LLMs) popularized by the release of GPT-3 and ChatGPT, LLMs have shown remarkable promise in programming-related tasks.
Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , Pierre Isabelle, Eugene Charniak, and Dekang Lin (Eds.). Association for Computational Linguistics, Philadelphia, Pennsylvania, USA, 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2009
Earlier work this paper cites.
Review of recent systems for automatic assessment of programming assignments. In Proceedings of the 10th Koli Calling International Conference on Computing Education Research (Koli, Finland) (Koli Calling ’10) . Association for Computing Machinery, New York, NY, USA, 86–93
Petri Ihantola, Tuukka Ahoniemi, Ville Karavirta, and Otto Seppälä. 2010 · 2010
Earlier work this paper cites.
Where Has The Time Gone? Faculty Activities and Time Commitments in the Online Classroom
B. Mandernach, Swinton Hudson, and Shanna Wise. 2013 · 2013
Earlier work this paper cites.
Automated Assessment of Programming Assignments. In Proceedings of the 3rd Computer Science Education Research Conference on Computer Science Education Research (Arnhem, Netherlands) (CSERC ’13) . Open Universiteit, Heerlen, Heerlen, NLD, 45–56
Vreda Pieterse. 2013 · 2013
Earlier work this paper cites.
Towards a Systematic Review of Automated Feedback Generation for Programming Exercises. In Proceedings of the 2016 ACM Conference on Innovation and Technology in Computer Science Education (Arequipa, Peru) (ITiCSE ’16) . Association for Computing Machinery, New York, NY, USA, 41–46
Hieke Keuning, Johan Jeuring, and Bastiaan Heeren. 2016 · 2016
Earlier work this paper cites.
Coderunner: a tool for assessing computer programming skills
Richard Lobb and Jenny Harlow. 2016 · 2016
Earlier work this paper cites.
On Novices’ Interaction with Compiler Error Messages: A Human Factors Approach. In Proceedings of the 2017 ACM Conference on International Computing Education Research (Tacoma, Washington, USA) (ICER ’17) . Association for Computing Machinery, New York, NY, USA, 74–82
James Prather, Raymond Pettit, Kayla Holcomb McMurry, Alani Peters, John Homer, Nevan Simone, and Maxine Cohen. 2017 · 2017
Earlier work this paper cites.
Application of Rubrics in the Classroom: A Vital Tool for Improvement in Assessment, Feedback and Learning
Faieza Chowdhury. 2018 · 2018
Earlier work this paper cites.
A Systematic Literature Review of Automated Feedback Generation for Programming Exercises
Hieke Keuning, Johan Jeuring, and Bastiaan Heeren. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
How Automated Feedback is Delivered Matters: Formative Feedback and Knowledge Transfer. In 2019 IEEE Frontiers in Education Conference (FIE) . 1–6
Qiang Hao and Michail Tsikerdekis. 2019 · 2019
Earlier work this paper cites.
Automatic Grading of Programming Assignments: An Approach Based on Formal Semantics. In 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engineering Education and Training (ICSE-SEET) . 126–137
Xiao Liu, Shuai Wang, Pei Wang, and Dinghao Wu. 2019 · 2019
Earlier work this paper cites.
Evaluation Rubric for Computational Thinking Concepts. In 2019 IEEE 19th International Conference on Advanced Learning Technologies (ICALT) , Vol. 2161-377X. 279–281
Christiano Otero Avila, Luciana Foss, Adriana Bordini, Maria Simone Debacco, and Simone André da Costa Cavalheiro. 2019 · 2019
Earlier work this paper cites.
Automatic Generation of Programming Exercises and Code Explanations Using Large Language Models. In Proceedings of the 2022 ACM Conference on International Computing Education Research - Volume 1 (ICER 2022) . ACM, 27–43
Sami Sarsa, Paul Denny, Arto Hellas, and Juho Leinonen. 2022 · 2022
Earlier work this paper cites.
Expectation vs Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI EA ’22) . Association for Computing Machinery, New York, NY, USA, Article 332, 7 pages
Priyan Vaithilingam, Tianyi Zhang, and Elena L. Glassman. 2022 · 2022
Cited alongside, same era.
Exploring the Responses of Large Language Models to Beginner Programmers’ Help Requests. In Proceedings of the 2023 ACM Conference on International Computing Education Research V.1 (ICER 2023) . ACM, 93–105
Arto Hellas, Juho Leinonen, Sami Sarsa, Charles Koutcheme, Lilja Kujanpää, and Juha Sorva. 2023 · 2023
Cited alongside, same era.
AI-TA: Towards an Intelligent Question-Answer Teaching Assistant using Open-Source LLMs
Yann Hicke, Anmol Agarwal, Qianou Ma, and Paul Denny. 2023 · 2023
Cited alongside, same era.
Feedback-Generation for Programming Exercises With GPT-4. In Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2024) . ACM, 31–37
Imen Azaiz, Natalie Kiesler, and Sven Strickroth. 2024 · 2024
Later among the works it cites.
Experiences from Integrating Large Language Model Chatbots into the Classroom
Arto Hellas, Juho Leinonen, and Leo Leppänen. 2024 · 2024
Later among the works it cites.
Evaluating the Application of Large Language Models to Generate Feedback in Programming Education. In 2024 IEEE Global Engineering Education Conference (EDUCON) . IEEE, 1–5
Sven Jacobs and Steffen Jaschke. 2024 · 2024
Later among the works it cites.
ChatGPT in the Classroom: An Analysis of Its Strengths and Weaknesses for Solving Undergraduate Computer Science Questions. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (Portland, OR, USA) (SIGCSE 2024) . Association for Computing Machinery, New York, NY, USA, 625–631
Ishika Joshi, Ritvik Budhiraja, Harshal Dev, Jahnvi Kadia, Mohammad Osama Ataullah, Sayan Mitra, Harshal D. Akolekar, and Dhruv Kumar. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ishika Joshi, Ritvik Budhiraja, Pranav Deepak Tanna, Lovenya Jain, Mihika Deshpande, Arjun Srivastava, Srinivas Rallapalli, Harshal D Akolekar, Jagat Sesh Challa, and Dhruv Kumar. 2023 · 2023
Cited alongside, same era.
Exploring the Potential of Large Language Models to Generate Formative Programming Feedback
Natalie Kiesler, Dominic Lohr, and Hieke Keuning. 2023 · 2023
Cited alongside, same era.
CodeHelp: Using Large Language Models with Guardrails for Scalable Support in Programming Classes
Mark Liffiton, Brad Sheese, Jaromir Savelka, and Paul Denny. 2023 · 2023
Cited alongside, same era.
Reduced grading in assessment: A scoping review
Dan-Anders Normann, Lise Vikan Sandvik, and Henning Fjørtoft. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2023
Cited alongside, same era.
Large Language Models (GPT) for automating feedback on programming assignments
Maciej Pankiewicz and Ryan S. Baker. 2023 · 2023
Cited alongside, same era.
Generative AI for Programming Education: Benchmarking ChatGPT, GPT-4, and Human Tutors
Tung Phung, Victor-Alexandru Pădurean, José Cambronero, Sumit Gulwani, Tobias Kohn, Rupak Majumdar, Adish Singla, and Gustavo Soares. 2023 · 2023
Cited alongside, same era.
CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code
Shuyan Zhou, Uri Alon, Sumit Agarwal, and Graham Neubig. 2023 · 2023
Cited alongside, same era.
Pingouin
2024 · 2024
Cited alongside, same era.
Later among the works it cites.
Automated Grading and Feedback Tools for Programming Education: A Systematic Review
Marcus Messer, Neil C. C. Brown, Michael Kölling, and Miaojing Shi. 2024 · 2024
Later among the works it cites.
Large Language Models in Computer Science Education: A Systematic Literature Review
Nishat Raihan, Mohammed Latif Siddiq, Joanna C. S. Santos, and Marcos Zampieri. 2024 · 2024
Later among the works it cites.
CodeJudge: Evaluating Code Generation with Large Language Models
Weixi Tong and Tianyi Zhang. 2024 · 2024
Later among the works it cites.
CodEv: An Automated Grading Framework Leveraging Large Language Models for Consistent and Constructive Feedback. In 2024 IEEE International Conference on Big Data (BigData) . IEEE, 5442–5449
En-Qi Tseng, Pei-Cing Huang, Chan Hsu, Peng-Yi Wu, Chan-Tung Ku, and Yihuang Kang. 2024 · 2024
Later among the works it cites.
Grade Like a Human: Rethinking Automated Assessment with Large Language Models
Wenjing Xie, Juxin Niu, Chun Jason Xue, and Nan Guan. 2024 · 2024
Later among the works it cites.
BeGrading: large language models for enhanced feedback in programming education
Mina Yousef, Kareem Mohamed, Walaa Medhat, Ensaf Hussein Mohamed, Ghada Khoriba, and Tamer Arafa. 2024 · 2024
Later among the works it cites.
ICE-Score: Instructing Large Language Models to Evaluate Code
Terry Yue Zhuo. 2024 · 2024
Later among the works it cites.
SedarEval: Automated Evaluation using Self-Adaptive Rubrics
Zhiyuan Fan, Weinong Wang, Xing Wu, and Debing Zhang. 2025 · 2025
Closest in time.
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation
Nachiket Kotalwar, Alkis Gotovos, and Adish Singla. 2025 · 2025
Closest in time.
Evaluating Language Models for Generating and Judging Programming Feedback. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 1 (Pittsburgh, PA, USA) (SIGCSETS 2025) . Association for Computing Machinery, New York, NY, USA, 624–630
Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen, Syed Ashraf, and Paul Denny. 2025 · 2025
Closest in time.
Large Language Models as Evaluators in Education: Verification of Feedback Consistency and Accuracy
Hyein Seo, Taewook Hwang, Jeesu Jung, Hyeonseok Kang, Hyuk Namgoong, Yohan Lee, and Sangkeun Jung. 2025 · 2025
Closest in time.