Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable performance on assisting humans in programming and facilitating programming automation.
Profile guided code positioning
Karl Pettis and Robert C. Hansen. 1990 · 1990
Earlier work this paper cites.
Refactoring - Improving the Design of Existing Code
Martin Fowler. 1999 · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Chin-Yew Lin and Eduard H. Hovy. 2003 · 2003
Earlier work this paper cites.
Code generation in the polyhedral model is easier than you think
Cédric Bastoul. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Measurement and quality in object-oriented design
Radu Marinescu. 2005 · 2005
Earlier work this paper cites.
A metric-based heuristic framework to detect object-oriented design flaws
Mazeiar Salehie, Shimin Li, and Ladan Tahvildari. 2006 · 2006
Earlier work this paper cites.
A practical automatic polyhedral parallelizer and locality optimizer
Uday Bondhugula, Albert Hartono, J. Ramanujam, and P. Sadayappan. 2008 · 2008
Earlier work this paper cites.
Chill: A framework for composing high-level loop transformations
Chun Chen, Jacqueline Chame, and Mary Hall. 2008 · 2008
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2009
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2009
Earlier work this paper cites.
Automatically finding patches using genetic programming
Westley Weimer, ThanhVu Nguyen, Claire Le Goues, and Stephanie Forrest. 2009 · 2009
Earlier work this paper cites.
On the use of automated text summarization techniques for summarizing source code
Sonia Haiduc, Jairo Aponte, Laura Moreno, and Andrian Marcus. 2010 · 2010
Earlier work this paper cites.
DECOR: A method for the specification and detection of code and design smells
Naouel Moha, Yann-Gaël Guéhéneuc, Laurence Duchien, and Anne-Françoise Le Meur. 2010 · 2010
Earlier work this paper cites.
Towards automatically generating summary comments for java methods
Giriprasad Sridhara, Emily Hill, Divya Muppaneni, Lori L. Pollock, and K. Vijay-Shanker. 2010 · 2010
Earlier work this paper cites.
Lexical statistical machine translation for language migration
Anh Tuan Nguyen, Tung Thanh Nguyen, and Tien N. Nguyen. 2013a · 2013
Earlier work this paper cites.
Semfix: program repair via semantic analysis
Hoang Duong Thien Nguyen, Dawei Qi, Abhik Roychoudhury, and Satish Chandra. 2013b · 2013
Earlier work this paper cites.
An automatic tool for tuning compiler optimizations
Dmitry Plotnikov, Dmitry Melnik, Mamikon Vardanyan, Ruben Buchatskiy, and Roman Zhuykov. 2013 · 2013
Earlier work this paper cites.
Defects4j: a database of existing faults to enable controlled testing studies for java programs
René Just, Darioush Jalali, and Michael D. Ernst. 2014 · 2014
Earlier work this paper cites.
The impact of code review coverage and code review participation on software quality: a case study of the qt, vtk, and ITK projects
Shane McIntosh, Yasutaka Kamei, Bram Adams, and Ahmed E. Hassan. 2014 · 2014
Earlier work this paper cites.
The strength of random search on automated program repair
Yuhua Qi, Xiaoguang Mao, Yan Lei, Ziying Dai, and Chengsong Wang. 2014 · 2014
Earlier work this paper cites.
Staged program repair with condition synthesis
Fan Long and Martin C. Rinard. 2015 · 2015
Earlier work this paper cites.
Autofdo: automatic feedback-directed optimization for warehouse-scale applications
Dehao Chen, David Xinliang Li, and Tipp Moseley. 2016 · 2016
Earlier work this paper cites.
Summarizing source code using a neural attention model
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016 · 2016
Earlier work this paper cites.
Latent predictor networks for code generation
Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomás Kociský, Fumin Wang, and Andrew W. Senior. 2016 · 2016
Earlier work this paper cites.
Designite: a software design quality assessment tool
Tushar Sharma, Pratibha Mishra, and Rohit Tiwari. 2016 · 2016
Earlier work this paper cites.
Deepcoder: Learning to write programs
Matej Balog, Alexander L. Gaunt, Marc Brockschmidt, Sebastian Nowozin, and Daniel Tarlow. 2017 · 2017
Cited alongside, same era.
Automatic configuration of GCC using irace
Leslie Pérez Cáceres, Federico Pagnozzi, Alberto Franzin, and Thomas Stützle. 2017 · 2017
Cited alongside, same era.
Robustfill: Neural program learning under noisy I/O
Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, and Pushmeet Kohli. 2017 · 2017
Cited alongside, same era.
Deepfix: Fixing common c language errors by deep learning
Rahul Gupta, Soham Pal, Aditya Kanade, and Shirish Shevade. 2017 · 2017
Cited alongside, same era.
Quixbugs: a multi-lingual program repair benchmark set based on the quixey challenge
Derrick Lin, James Koppel, Angela Chen, and Armando Solar-Lezama. 2017 · 2017
Cited alongside, same era.
Piecewise holistic autotuning of parallel programs with CERE
Mihail Popov, Chadi Akel, Yohan Chatelain, William Jalby, and Pablo de Oliveira Castro. 2017 · 2017
Using pre-trained models to boost code review automation
Rosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella, Denys Poshyvanyk, and Gabriele Bavota. 2022 · 2022
Later among the works it cites.
Whygen: Explaining ml-powered code generation by referring to training examples
Weixiang Yan and Yuanchun Li. 2022 · 2022
Later among the works it cites.
Xlcost: A benchmark dataset for cross-lingual code intelligence
Ming Zhu, Aneesh Jain, Karthik Suresh, Roshan Ravindran, Sindhu Tipirneni, and Chandan K. Reddy. 2022 · 2022
Later among the works it cites.
Palm 2 technical report
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Ábrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, and et al. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A syntactic neural model for general-purpose code generation
Pengcheng Yin and Graham Neubig. 2017 · 2017
Cited alongside, same era.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019 · 2019
Cited alongside, same era.
An empirical study on learning bug-fixing patches in the wild via neural machine translation
Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. 2019 · 2019
Cited alongside, same era.
Codemason: Binary-level profile-guided optimization
David Williams-King and Junfeng Yang. 2019 · 2019
Cited alongside, same era.
Improved code summarization via a graph neural network
Alexander LeClair, Sakib Haque, Lingfei Wu, and Collin McMillan. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
Multi-lingual evaluation of code generation models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li, Yuchen Tian, Ming Tan, Wasi Uddin Ahmad, Shiqi Wang, Qing Sun, Mingyue Shang, Sujan Kumar Gonugondla, Hantian Ding, Varun Kumar, Nathan Fulton, Arash Farahani, Siddhartha Jain, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, and Ramesh Nallapati. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023 · 2023
Closest in time.
Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou. 2023 · 2023
Closest in time.
Chatgpt outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Closest in time.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023 · 2023
Closest in time.
Mohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Long Do, Weishi Wang, Md. Rizwan Parvez, and Shafiq R. Joty. 2023 · 2023
Closest in time.
DS-1000: A natural and reliable benchmark for data science code generation
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-Tau Yih, Daniel Fried, Sida I. Wang, and Tao Yu. 2023 · 2023
Closest in time.
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023 · 2023
Closest in time.
Detecting code smells using industry-relevant data
Lech Madeyski and Tomasz Lewowski. 2023 · 2023
Closest in time.
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023 · 2023
Closest in time.
GPT-4 technical report
OpenAI. 2023 · 2023
Closest in time.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023 · 2023
Closest in time.
Exploring the effectiveness of large language models in generating unit tests
Mohammed Latif Siddiq, Joanna C. S. Santos, Ridwanul Hasan Tanvir, Noshin Ulfat, Fahmid Al Rifat, and Vinicius Carvalho Lopes. 2023 · 2023
Closest in time.
Towards a systematic approach to manual annotation of code smells
Jelena Slivka, Nikola Luburic, Simona Prokic, Katarina-Glorija Grujic, Aleksandar Kovacevic, Goran Sladic, and Dragan Vidakovic. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Closest in time.
Chatunitest: a chatgpt-based automated unit test generation tool
Zhuokui Xie, Yinghao Chen, Chen Zhi, Shuiguang Deng, and Jianwei Yin. 2023 · 2023
Closest in time.
Codetransocean: A comprehensive multilingual benchmark for code translation
Weixiang Yan, Yuchen Tian, Yunzhe Li, Qian Chen, and Wen Wang. 2023 · 2023
Closest in time.
Codereval: A benchmark of pragmatic code generation with generative pre-trained models
Hao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang, Qi Zhang, Yuchi Ma, Guangtai Liang, Ying Li, Tao Xie, and Qianxiang Wang. 2023 · 2023
Closest in time.
No more manual tests? evaluating and improving chatgpt for unit test generation
Zhiqiang Yuan, Yiling Lou, Mingwei Liu, Shiji Ding, Kaixin Wang, Yixuan Chen, and Xin Peng. 2023 · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023 · 2023
Closest in time.
Language agent tree search unifies reasoning acting and planning in language models
Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. 2023 · 2023
Closest in time.
Detecting code smells using deep learning
Ananta Kumar Das, Shikhar Yadav, and Subhasish Dhal. 2019 · 2086
Closest in time.