Fetching the paper…
Reading the bibliography…
LLM-based code generation tools are essential to help developers in the software development process.
The generalization of ‘student’s’ problem when several different population varlances are involved
Welch Bl · 1947
Earlier work this paper cites.
An application of hierarchical kappa-type statistics in the assessment of majority agreement among multiple observers
J Richard Landis and Gary G. Koch · 1977
Earlier work this paper cites.
The structure of modular program
Joshua Turner · 1980
Earlier work this paper cites.
Elements of survey sampling
Ravindra Pal Singh and Naurang Singh Mangat · 1996
Earlier work this paper cites.
Software development on internet time
Michael A Cusumano and David B Yoffie · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Spiral: Code generation for dsp transforms
Markus Puschel, José MF Moura, Jeremy R Johnson, David Padua, Manuela M Veloso, Bryan W Singer, Jianxin Xiong, Franz Franchetti, Aca Gacic, Yevgen Voronenko, et al · 2005
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
An association rule mining method for estimating the impact of project management policies on software quality, development time and effort
María N Moreno García, Isabel Ramos Román, Francisco J García Peñalvo, and Miguel Toro Bonilla · 2008
Earlier work this paper cites.
Logic codes generation and transmission using an encoding-decoding system
D Gifany, IS Amiri, M Ranjbar, and J Ali · 2013
Earlier work this paper cites.
Spoon: A Library for Implementing Analyses and Transformations of Java Source Code
Renaud Pawlak, Martin Monperrus, Nicolas Petitprez, Carlos Noguera, and Lionel Seinturier · 2015
Earlier work this paper cites.
Retrieval-based neural code generation
Shirley Anugrah Hayati, Raphaël Olivier, Pravalika Avvaru, Pengcheng Yin, Anthony Tomasic, and Graham Neubig · 2018
Earlier work this paper cites.
A novel neural source code representation based on abstract syntax tree
Jian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun, Kaixuan Wang, and Xudong Liu · 2019
Earlier work this paper cites.
Intellicode compose: Code generation using transformer
Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan · 2020
Earlier work this paper cites.
Treegen: A tree-based transformer architecture for code generation
Zeyu Sun, Qihao Zhu, Yingfei Xiong, Yican Sun, Lili Mou, and Lu Zhang · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
”github copilot. your ai pair programmer”
Available:https://copilot.github.com/ · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Honeysuckle: Annotation-guided code generation of in-app privacy notices
Tianshi Li, Elijah B Neundorfer, Yuvraj Agarwal, and Jason I Hong · 2021
Earlier work this paper cites.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu · 2021
Earlier work this paper cites.
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi · 2021
Earlier work this paper cites.
Empirical analysis of security vulnerabilities in python packages
Mahmoud Alfadel, Diego Elias Costa, and Emad Shihab · 2021
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston · 2021
Earlier work this paper cites.
Recent advances in intelligent source code generation: A survey on natural language based studies
Chen Yang, Yan Liu, and Changqing Yin · 2021
Earlier work this paper cites.
A survey of automatic code generation fromnatural language
Jiho Shin and Jaechang Nam · 2021
Earlier work this paper cites.
Generating code with the help of retrieved template functions and stack overflow answers
Dawn Drain, Changran Hu, Chen Wu, Mikhail Breslav, and Neel Sundaresan · 2021
Cited alongside, same era.
Retrieval augmented code generation and summarization
Md. Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang · 2021
Cited alongside, same era.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Cited alongside, same era.
Codet: Code generation with generated tests
Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen · 2022
Cited alongside, same era.
Cert: Continual pre-training on sketches for library-oriented code generation
Gpt-4 technical report
OpenAI · 2023
Closest in time.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Huai hsin Chi, Nathanael Scharli, and Denny Zhou · 2023
Closest in time.
Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou · 2023
Closest in time.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2023
Closest in time.
Better context makes better code language models: A case study on function call argument completion
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daoguang Zan, Bei Chen, Dejian Yang, Zeqi Lin, Minsu Kim, Bei Guan, Yongji Wang, Weizhu Chen, and Jian-Guang Lou · 2022
Cited alongside, same era.
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Haiquan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2022
Cited alongside, same era.
Incoder: A generative model for code infilling and synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida I. Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen tau Yih, Luke Zettlemoyer, and Mike Lewis · 2022
Cited alongside, same era.
Syntax-aware on-the-fly code completion
Wannita Takerngsaksiri, Chakkrit Kla Tantithamthavorn, and Yuankui Li · 2022
Cited alongside, same era.
A systematic evaluation of large language models of code
Frank F. Xu, Uri Alon, Graham Neubig, and Vincent J. Hellendoorn · 2022
Cited alongside, same era.
Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models
Priyan Vaithilingam, Tianyi Zhang, and Elena L Glassman · 2022
Cited alongside, same era.
Cocomic: Code completion by jointly modeling in-file and cross-file context
Yangruibo Ding, Zijian Wang, Wasi Uddin Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang · 2022
Cited alongside, same era.
Proceedings of the 2022 national toxicology program satellite symposium
Erin M. Quist, Shambhunath Choudhary, Richard Lang, Debra A. Tokarz, Mark James Hoenerhoff, Jonathan Nagel, and Jeffrey I. Everitt · 2022
Cited alongside, same era.
Hengzhi Pei, Jinman Zhao, Leonard Lausen, Sheng Zha, and George Karypis · 2023
Closest in time.
Repofusion: Training code models to understand your repository
Disha Shrivastava, Denis Kocetkov, Harm de Vries, Dzmitry Bahdanau, and Torsten Scholak · 2023
Closest in time.
In-context instruction learning
Seonghyeon Ye, Hyeonbin Hwang, Sohee Yang, Hyeongu Yun, Yireun Kim, and Minjoon Seo · 2023
Closest in time.
Gpt-3.5-turbo-16k
OpenAI · 2023
Closest in time.
Api entity and relation joint extraction from text via dynamic prompt-tuned language model
Qing Huang, Yanbang Sun, Zhenchang Xing, Mingming Yu, Xiwei Xu, and Qinghua Lu · 2023
Closest in time.
Codegen4libs: A two-stage approach for library-oriented code generation
Mingwei Liu, Tianyong Yang, Yiling Lou, Xueying Du, Ying Wang, and Xin Peng · 2023
Closest in time.
From misuse to mastery: Enhancing code generation with knowledge-driven ai chaining
Xiaoxue Ren, Xinyuan Ye, Dehai Zhao, Zhenchang Xing, and Xiaohu Yang · 2023
Closest in time.
Faithful chain-of-thought reasoning
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch · 2023
Closest in time.
https://platform.openai.com/docs/guides/embeddings/what-are-embeddings , note = Accessed: 2023.9, 2023
OpenAI · 2023
Closest in time.
Vec2vec: A compact neural network approach for transforming text embeddings with high fidelity
Andrew Gao · 2023
Closest in time.
Jina embeddings 2: 8192-token general-purpose text embeddings for long documents
Michael Günther, Jackmin Ong, Isabelle Mohr, Alaeddine Abdessalem, Tanguy Abel, Mohammad Kalim Akram, Susana Guzman, Georgios Mastrapas, Saba Sturua, Bo Wang, Maximilian Werk, Nan Wang, and Han Xiao · 2023
Closest in time.
Evaluating instruction-tuned large language models on code comprehension and generation
Zhiqiang Yuan, Junwei Liu, Qiancheng Zi, Mingwei Liu, Xin Peng, and Yiling Lou · 2023
Closest in time.
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang · 2023
Closest in time.
Llm is like a box of chocolates: the non-determinism of chatgpt in code generation
Shuyin Ouyang, Jie M Zhang, Mark Harman, and Meng Wang · 2023
Closest in time.
Unveiling memorization in code models
Zhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi, Dongsun Kim, Donggyun Han, and David Lo · 2023
Closest in time.
https://github.com/gak/pycallgraph , note = Accessed: 2023.9, 2023
Pycallgraph · 2023
Closest in time.
https://github.com/davidfraser/pyan , note = Accessed: 2023.9, 2023
Pyan · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Closest in time.
How language model hallucinations can snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A Smith · 2023
Closest in time.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Closest in time.
Repocoder: Repository-level code completion through iterative retrieval and generation
Fengji Zhang, B. Chen, Yue Zhang, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen · 2023
Closest in time.
Compositional api recommendation for library-oriented code generation
Zexiong Ma, Shengnan An, Bing Xie, and Zeqi Lin · 2024
Closest in time.
On leakage of code generation evaluation datasets
Alexandre Matton, Tom Sherborne, Dennis Aumiller, Elena Tommasone, Milad Alizadeh, Jingyi He, Raymond Ma, Maxime Voisin, Ellen Gilsenan-McMahon, and Matthias Gall’e · 2024
Closest in time.
A solution toward transparent and practical ai regulation: Privacy nutrition labels for open-source generative ai-based applications
Meixue Si, Shidong Pan, Dianshu Liao, Xiaoyu Sun, Zhenyuan Tao, Wenchang Shi, and Zhenchang Xing · 2024
Closest in time.