Fetching the paper…
Reading the bibliography…
Large language models have demonstrated great potential to assist programmers in generating code.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Classifier technology and the illusion of progress
David J. Hand. 2006 · 2006
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2009
Earlier work this paper cites.
Lexical statistical machine translation for language migration
Anh Tuan Nguyen, Tung Thanh Nguyen, and Tien N Nguyen. 2013 · 2013
Earlier work this paper cites.
Learning natural coding conventions
Miltiadis Allamanis, Earl T Barr, Christian Bird, and Charles Sutton. 2014 · 2014
Earlier work this paper cites.
Phrase-based statistical translation of programming languages
Svetoslav Karaivanov, Veselin Raychev, and Martin Vechev. 2014 · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems
Martín Abadi, Ashish Agarwal, and Paul Barham et al. 2015 · 2015
Earlier work this paper cites.
Will they like this? evaluating code contributions with language models
Vincent J Hellendoorn, Premkumar T Devanbu, and Alberto Bacchelli. 2015 · 2015
Earlier work this paper cites.
Antonio Valerio Miceli Barone and Rico Sennrich. 2017 · 2017
Earlier work this paper cites.
A survey of machine learning for big code and naturalness
Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton. 2018 · 2018
Earlier work this paper cites.
The adverse effects of code duplication in machine learning models of code
Miltiadis Allamanis. 2019 · 2019
Earlier work this paper cites.
When code completion fails: A case study on real-world completions
Vincent J. Hellendoorn, Sebastian Proksch, Harald C. Gall, and Alberto Bacchelli. 2019 · 2019
Earlier work this paper cites.
Spoc: Search-based pseudocode to code
Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, and Percy S Liang. 2019 · 2019
Cited alongside, same era.
Beyond accuracy: Grounding evaluation metrics for human-machine learning systems
Praveen Chandar, Fernando Diaz, and Brian St. Thomas. 2020 · 2020
Cited alongside, same era.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Cited alongside, same era.
Unsupervised translation of programming languages
Baptiste Roziere, Marie-Anne Lachaux, Lowik Chanussot, and Guillaume Lample. 2020 · 2020
Cited alongside, same era.
Intellicode compose: Code generation using transformer
Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020 · 2020
Cited alongside, same era.
Unified pre-training for program understanding and generation
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Later among the works it cites.
Perfection not required? human-ai partnerships in code translation
Justin D Weisz, Michael Muller, Stephanie Houde, John Richards, Steven I Ross, Fernando Martinez, Mayank Agarwal, and Kartik Talamadupula. 2021 · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022 · 2022
Closest in time.
Incoder: A generative model for code infilling and synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen-tau Yih, Luke Zettlemoyer, and Mike Lewis. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Cited alongside, same era.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Cited alongside, same era.
The space of developer productivity: There’s more to it than you think
Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Tom Zimmermann, Brian Houck, and Jenna Butler. 2021 · 2021
Cited alongside, same era.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al. 2021 · 2021
Cited alongside, same era.
Measurement and fairness
Abigail Z. Jacobs and Hanna Wallach. 2021 · 2021
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021 · 2021
Cited alongside, same era.
Quantifying GitHub Copilot’s impact on developer productivity and happiness
Eirini Kalliamvakou. 2022 · 2022
Closest in time.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022 · 2022
Closest in time.
A conversational paradigm for program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022 · 2022
Closest in time.
The fallacy of AI functionality
Inioluwa Deborah Raji, I. Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst. 2022 · 2022
Closest in time.
Reliance on metrics is a fundamental challenge for ai
Rachel L. Thomas and David Uminsky. 2022 · 2022
Closest in time.
A systematic evaluation of large language models of code
Frank F Xu, Uri Alon, Graham Neubig, and Vincent J Hellendoorn. 2022 · 2022
Closest in time.
Deconstructing nlg evaluation: Evaluation practices, assumptions, and their implications
Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, and Alexandra Olteanu. 2022 · 2022
Closest in time.
Productivity assessment of neural code completion
Albert Ziegler, Eirini Kalliamvakou, X. Alice Li, Andrew Rice, Devon Rifkin, Shawn Simister, Ganesh Sittampalam, and Edward Aftandilian. 2022 · 2022
Closest in time.