Fetching the paper…
Reading the bibliography…
Ensemble learning has been widely used in machine learning to improve model robustness, accuracy, and generalization, but has not yet been applied to code generation tasks with large language models (LLMs).
Stacked generalization
David H Wolpert · 1992
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I. Jordan and Robert A. Jacobs · 1994
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
On combining classifiers
Josef Kittler, Mohamad Hatef, Robert P W Duin, and Jiri Matas · 1998
Earlier work this paper cites.
Differential static analysis: opportunities, applications, and challenges
Shuvendu K Lahiri, Kapil Vaswani, and C AR Hoare · 2010
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
https://github.com/pschanely/CrossHair, 2020
Crosshair: Symbolic execution for python · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Hydiff: Hybrid differential software analysis
Yannic Noller, Corina S. Păsăreanu, Marcel Böhme, Youcheng Sun, Hoang Lam Nguyen, and Lars Grunske · 2020
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Shiqi Ren, Zhiyang Tu, Deng Cai, Jun Zhang, and Yanzhi Chen · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
https://codellama.dev/about, 2022
Codellama · 2022
Earlier work this paper cites.
https://deepseekcoder.github.io/t, 2022
Deepseekcoder · 2022
Earlier work this paper cites.
Codet: Code generation with generated tests
Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Prompt-tuned code language model as a neural knowledge base for type inference in statically-typed partial code
Qing Huang, Zhiqiang Yuan, Zhenchang Xing, Xiwei Xu, Liming Zhu, and Qinghua Lu · 2022
Cited alongside, same era.
Is github copilot a substitute for human pair-programming? an empirical study
Saki Imai · 2022
Cited alongside, same era.
Jigsaw: Large language models meet program synthesis
Naman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan, Suresh Parthasarathy, Sriram Rajamani, and Rahul Sharma · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Cited alongside, same era.
Large language models for software engineering: Survey and open problems
Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M Zhang · 2023
Later among the works it cites.
Invalidator: Automated patch correctness assessment via semantic and syntactic reasoning
Thanh Le-Cong, Duc-Minh Luong, Xuan Bach D Le, David Lo, Nhat-Hoa Tran, Bui Quang-Huy, and Quyet-Thang Huynh · 2023
Later among the works it cites.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang · 2023
Later among the works it cites.
Fill in the blank: Context-aware automated text input generation for mobile gui testing
Zhe Liu, Chunyang Chen, Junjie Wang, Xing Che, Yuekai Huang, Jun Hu, and Qing Wang · 2023
Later among the works it cites.
Automated program repair based on code review: How do pre-trained transformer models perform?
Rishov Paul, Md Mohib Hossain, Masum Hasan, and Anindya Iqbal · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey of ensemble learning: Concepts, algorithms, applications, and prospects
Ibomoiye Domor Mienye and Yanxia Sun · 2022
Cited alongside, same era.
A systematic evaluation of large language models of code
Frank F Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn · 2022
Cited alongside, same era.
Productivity assessment of neural code completion
Albert Ziegler, Eirini Kalliamvakou, X Alice Li, Andrew Rice, Devon Rifkin, Shawn Simister, Ganesh Sittampalam, and Edward Aftandilian · 2022
Cited alongside, same era.
https://chat.openai.com/, 2023
Chatgpt · 2023
Cited alongside, same era.
https://github.com/features/copilot, 2023
Github copilot · 2023
Cited alongside, same era.
https://openai.com/gpt-4, 2023
Gpt-4 · 2023
Cited alongside, same era.
https://openai.com/blog/openai-codex, 2023
Openai codex · 2023
Cited alongside, same era.
Examining zero-shot vulnerability repair with large language models
Hammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri, and Brendan Dolan-Gavitt · 2023
Later among the works it cites.
The best of both worlds: Combining learned embeddings with engineered features for accurate prediction of correct patches
Haoye Tian, Kui Liu, Yinghua Li, Abdoul Kader Kaboré, Anil Koyuncu, Andrew Habib, Li Li, Junhao Wen, Jacques Klein, and Tegawendé F Bissyandé · 2023
Later among the works it cites.
Automated program repair in the era of large pre-trained language models
Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang · 2023
Later among the works it cites.
Keep the conversation going: Fixing 162 out of 337 bugs for $0.42 each using chatgpt
Chunqiu Steven Xia and Lingming Zhang · 2023
Later among the works it cites.
Livecodebench: Benchmarking large language models for code in the wild
Frank F Xu, Joon Sung Shin, Gabriel Poesia, Yao Luan, Pengcheng Yin, Xi Victoria Lin, Yi-Ting Chiu, Scott Lundberg, Hannaneh Hajishirzi, Mike Guo, et al · 2023
Later among the works it cites.
Codellama: Open foundation models for code
Mateusz Ziemniak, Johannes Meyer, Victor Bajona, Asya Glaese, Guillaume Belrose, Sebastian Borgeaud, Mayur Daswani, Luca Demontis, Marylou Gabrié, Felix Hill, et al · 2023
Later among the works it cites.
Llm-based test-driven interactive code generation: User study and empirical evaluation
Sarah Fakhoury, Aaditya Naik, Georgios Sakkas, Saikat Chakraborty, and Shuvendu K Lahiri · 2024
Later among the works it cites.
Large language models for software engineering: A systematic literature review
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang · 2024
Later among the works it cites.
From llms to llm-based agents for software engineering: A survey of current, challenges and future
Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan, Bo Li, and Huaming Chen · 2024
Later among the works it cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri · 2025
Closest in time.