Fetching the paper…
Reading the bibliography…
Large language models (LLMs), such as Codex, hold great promise in enhancing programming education by automatically generating feedback for students.
The χ \chi 2 Test of Goodness of Fit
William G Cochran · 1952
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
What Would Other Programmers Do: Suggesting Solutions to Error Messages
Björn Hartmann, Daniel MacDougall, Joel Brandt, and Scott R. Klemmer · 2010
Earlier work this paper cites.
Automated Feedback Generation for Introductory Programming Assignments
Rishabh Singh, Sumit Gulwani, and Armando Solar-Lezama · 2013
Earlier work this paper cites.
Five Ways to Look at Cohen’s Kappa
Matthijs J Warrens · 2015
Earlier work this paper cites.
An Effective Approach to Enhancing Compiler Error Messages
Brett A. Becker · 2016
Earlier work this paper cites.
Description2Code Dataset
Ethan Caballero and Ilya Sutskever · 2016
Earlier work this paper cites.
On Novices’ Interaction with Compiler Error Messages: A Human Factors Approach
James Prather, Raymond Pettit, Kayla Holcomb McMurry, Alani L. Peters, John Homer, Nevan Simone, and Maxine S. Cohen · 2017
Earlier work this paper cites.
Writing Reusable Code Feedback at Scale with Mixed-Initiative Program Synthesis
Andrew Head, Elena L. Glassman, Gustavo Soares, Ryo Suzuki, Lucas Figueredo, Loris D’Antoni, and Björn Hartmann · 2017
Earlier work this paper cites.
Automated Clustering and Program Repair for Introductory Programming Assignments
Sumit Gulwani, Ivan Radicek, and Florian Zuleger · 2018
Earlier work this paper cites.
Neuro-Symbolic Program Corrector for Introductory Programming Assignments
Sahil Bhatia, Pushmeet Kohli, and Rishabh Singh · 2018
Earlier work this paper cites.
Understanding Back-Translation at Scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier · 2018
Earlier work this paper cites.
Deep Reinforcement Learning for Syntactic Error Repair in Student Programs
Rahul Gupta, Aditya Kanade, and Shirish K. Shevade · 2019
Cited alongside, same era.
Toward Data-Driven Example Feedback for Novice Programming
Rui Zhi, Samiha Marwan, Yihuan Dong, Nicholas Lytle, Thomas W. Price, and Tiffany Barnes · 2019
Cited alongside, same era.
The Error Behind The Message: Finding the Cause of Error Messages in Python
Tobias Kohn · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners
Tom B. Brown and et al · 2020
Cited alongside, same era.
Tell Me What’s Wrong: A Python IDE with Error Messages
Tobias Kohn and Bill Z. Manaris · 2020
Cited alongside, same era.
Program Synthesis with Pragmatic Communication
Yewen Pu, Kevin Ellis, Marta Kryven, Josh Tenenbaum, and Armando Solar-Lezama · 2020
Cited alongside, same era.
I Speak, You Verify: Toward Trustworthy Neural Program Synthesis
Darren Key, Wen-Ding Li, and Kevin Ellis · 2022
Later among the works it cites.
Balancing Cost and Quality: An Exploration of Human-in-the-Loop Frameworks for Automated Short Answer Scoring
Hiroaki Funayama, Tasuku Sato, Yuichiroh Matsubayashi, Tomoya Mizumoto, Jun Suzuki, and Kentaro Inui · 2022
Later among the works it cites.
Adaptive Scaffolding in Block-Based Programming via Synthesizing New Tasks as Pop Quizzes
Ahana Ghosh, Sebastian Tschiatschek, Sam Devlin, and Adish Singla · 2022
Later among the works it cites.
Automated Repair of Programs from Large Language Models
Zhiyu Fan, Xiang Gao, Abhik Roychoudhury, and Shin Hwei Tan · 2022
Later among the works it cites.
Neurosymbolic Repair for Low-Code Formula Languages
Rohan Bavishi, Harshit Joshi, José Cambronero, Anna Fariha, Sumit Gulwani, Vu Le, Ivan Radicek, and Ashish Tiwari · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evaluating Large Language Models Trained on Code
Mark Chen and et al · 2021
Cited alongside, same era.
What Does Saying That ‘Programming is Hard’ Really Say, and About Whom?
Brett A. Becker · 2021
Cited alongside, same era.
The Robots Are Coming: Exploring the Implications of OpenAI Codex on Introductory Programming
James Finnie-Ansley, Paul Denny, Brett A. Becker, Andrew Luxton-Reilly, and James Prather · 2022
Cited alongside, same era.
Automatic Generation of Programming Exercises and Code Explanations Using Large Language Models
Sami Sarsa, Paul Denny, Arto Hellas, and Juho Leinonen · 2022
Cited alongside, same era.
Repairing Bugs in Python Assignments Using Large Language Models
Jialu Zhang, José Cambronero, Sumit Gulwani, Vu Le, Ruzica Piskac, Gustavo Soares, and Gust Verbruggen · 2022
Cited alongside, same era.
Codeforces
Mikhail Mirzayanov
Cited in the paper.
Competition-Level Code Generation with AlphaCode
Yujia Li and et al · 2022
Later among the works it cites.
Experiences from Using Code Explanations Generated by Large Language Models in a Web Software Development E-Book
Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny, Seth Bernstein, and Juho Leinonen · 2023
Closest in time.
Using Large Language Models to Enhance Programming Error Messages
Juho Leinonen, Arto Hellas, Sami Sarsa, Brent N. Reeves, Paul Denny, James Prather, and Brett A. Becker · 2023
Closest in time.
What is Your Biggest Pain Point? An Investigation of CS Instructor Obstacles, Workarounds, and Desires
Samim Mirhosseini, Austin Z. Henley, and Chris Parnin · 2023
Closest in time.
Repair is Nearly Generation: Multilingual Program Repair with LLMs
Harshit Joshi, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le, Ivan Radicek, and Gust Verbruggen · 2023
Closest in time.
The AI Teacher Test: Measuring the Pedagogical Ability of Blender and GPT-3 in Educational Dialogues
Anaïs Tack and Chris Piech · 2023
Closest in time.