Fetching the paper…
Reading the bibliography…
Over the past few years, Large Language Models of Code (Code LLMs) have started to have a significant impact on programming practice.
English as a Very High Level Language for Simulation Programming. In Proceedings of the ACM SIGPLAN Symposium on Very High Level Languages . Association for Computing Machinery, New York, NY, USA, 91–100
George E. Heidorn. 1974 · 1974
Earlier work this paper cites.
Automatic Evaluation of Machine Translation Quality Using Longest Common Subsequence and Skip-Bigram Statistics. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04) . Barcelona, Spain, 605–612
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Understanding the Factors That Impact the Popularity of GitHub Repositories. In 2016 IEEE International Conference on Software Maintenance and Evolution (ICSME) . 334–344
Hudson Borges, Andre Hora, and Marco Tulio Valente. 2016 · 2016
Earlier work this paper cites.
GitHub on BigQuery: Analyze All the Open Source Code
Felipe Hoffa. 2016 · 2016
Earlier work this paper cites.
Automated Unit Test Generation for Python. In Proceedings of the 12th Symposium on Search-based Software Engineering (SSBSE 2020, Bari, Italy, October 7–8) (Lecture Notes in Computer Science, Vol. 12420) . Springer, 9–24
Stephan Lukasczyk, Florian Kroiß, and Gordon Fraser. 2020 · 2020
Earlier work this paper cites.
ZeRO: Memory Optimizations toward Training Trillion Parameter Models. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , Qun Liu and David Schlangen (Eds.). Association for Computational Linguistics, 38–45
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Earlier work this paper cites.
Program Synthesis with Large Language Models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Leveraging Automated Unit Tests for Unsupervised Code Translation. In International Conference on Learning Representations
Baptiste Roziere, Jie Zhang, Francois Charton, Mark Harman, Gabriel Synnaeve, and Guillaume Lample. 2021 · 2021
Earlier work this paper cites.
Multilingual training for software engineering. In Proceedings of the 44th International Conference on Software Engineering (Pittsburgh, Pennsylvania) (ICSE ’22) . Association for Computing Machinery, New York, NY, USA, 1443–1455
Toufique Ahmed and Premkumar Devanbu. 2022 · 2022
Earlier work this paper cites.
Multi-Lingual Evaluation of Code Generation Models. In The Eleventh International Conference on Learning Representations
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li, Yuchen Tian, Ming Tan, Wasi Uddin Ahmad, Shiqi Wang, Qing Sun, Mingyue Shang, Sujan Kumar Gonugondla, Hantian Ding, Varun Kumar, Nathan Fulton, Arash Farahani, Siddhartha Jain, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Sudipta Sengupta, Dan Roth, and Bing Xiang. 2022 · 2022
Earlier work this paper cites.
Code Generation Tools (Almost) for Free? A Study of Few-Shot, Pre-Trained Language Models on Code
Patrick Bareiß, Beatriz Souza, Marcelo d’Amorim, and Michael Pradel. 2022 · 2022
Earlier work this paper cites.
On the Transferability of Pre-Trained Language Models for Low-Resource Programming Languages. In IEEE/ACM International Conference on Program Comprehension (ICPC) . Association for Computing Machinery, 401–412
Fuxiang Chen, Fatemeh H. Fard, David Lo, and Timofey Bryksin. 2022 · 2022
Earlier work this paper cites.
An empirical analysis of compute-optimal large language model training. In Advances in Neural Information Processing Systems , Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.)
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack William Rae, and Laurent Sifre. 2022 · 2022
Earlier work this paper cites.
Deduplicating Training Data Makes Language Models Better. In Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Dublin, Ireland, 8424–8445
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022 · 2022
Earlier work this paper cites.
Training Language Models to Follow Instructions with Human Feedback. In Advances in Neural Information Processing Systems (NeurIPS) , Vol. 35. Curran Associates, Inc
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
A Systematic Evaluation of Large Language Models of Code. In Deep Learning for Code Workshop (DL4C)
Frank F. Xu, Uri Alon, Graham Neubig, and Vincent J. Hellendoorn. 2022 · 2022
Earlier work this paper cites.
Productivity Assessment of Neural Code Completion. In Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming (San Diego, CA, USA) (MAPS 2022) . Association for Computing Machinery, New York, NY, USA, 21–29
Albert Ziegler, Eirini Kalliamvakou, X. Alice Li, Andrew Rice, Devon Rifkin, Shawn Simister, Ganesh Sittampalam, and Edward Aftandilian. 2022 · 2022
Earlier work this paper cites.
SantaCoder: Don’t Reach for the Stars!. In Deep Learning for Code Workshop (DL4C)
Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, Logesh Kumar Umapathi, Carolyn Jane Anderson, Yangtian Zi, Joel Lamy Poirier, Hailey Schoelkopf, Sergey Troshin, Dmitry Abulkhanov, Manuel Romero, Michael Lappert, Francesco De Toni, Bernardo García del Río, Qian Liu, Shamik Bose, Urvashi Bhattacharyya, Terry Yue Zhuo, Ian Yu, Paulo Villegas, Marco Zocca, Sourab Mangrulkar, David Lansky, Huu Nguyen, Danish Contractor, Luis Villa, Jia Li, Dzmitry Bahdanau, Yacine Jernite, Sean Hughes, Daniel Fried, Arjun Guha, Harm de Vries, and Leandro von Werra. 2023 · 2023
Earlier work this paper cites.
PaLM 2 Technical Report
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Music Li, Wei Li, YaGuang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. 2023 · 2023
Earlier work this paper cites.
Model Card and Evaluations for Claude Models
Anthropic. 2023a · 2023
Earlier work this paper cites.
Terms of Service
Anthropic. 2023b · 2023
Cited alongside, same era.
MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation
Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q. Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda. 2023 · 2023
Cited alongside, same era.
Code Alpaca: An Instruction-following LLaMA model for code generation
Sahil Chaudhary. 2023 · 2023
Cited alongside, same era.
Data Race Detection Using Large Language Models. In Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis (SC-W) . Association for Computing Machinery, New York, NY, USA, 215–223
Le Chen, Xianzhong Ding, Murali Emani, Tristan Vanderbruggen, Pei-Hung Lin, and Chunhua Liao. 2023 · 2023
Cited alongside, same era.
ML-powered Coding Companion – Amazon CodeWhisperer – Amazon Web Services
CodeWhisperer. 2023 · 2023
Cited alongside, same era.
Static Type Checker for Python
Pyright. 2023 · 2023
Closest in time.
Replit Code v1.3
Replit. 2023 · 2023
Closest in time.
The Programmer’s Assistant: Conversational Interaction with a Large Language Model for Software Development. In Proceedings of the 28th International Conference on Intelligent User Interfaces (Sydney, NSW, Australia) (IUI ’23) . Association for Computing Machinery, New York, NY, USA, 491–514
Steven I. Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D. Weisz. 2023 · 2023
Closest in time.
Code Llama: Open Foundation Models for Code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023 · 2023
Closest in time.
Adaptive Test Generation Using a Large Language Model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Github Copilot Your AI pair programmer
Github Copilot. 2023 · 2023
Cited alongside, same era.
Go smol or go home
Harm de Vries. 2023 · 2023
Cited alongside, same era.
Baldur: Whole-Proof Generation and Repair with Large Language Models. In ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE) (6–8). San Fransisco, CA, USA
Emily First, Markus Rabe, Talia Ringer, and Yuriy Brun. 2023 · 2023
Cited alongside, same era.
Generative AI Terms of Service
Google. 2023 · 2023
Cited alongside, same era.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023 · 2023
Cited alongside, same era.
Repair Is Nearly Generation: Multilingual Program Repair with LLMs
Harshit Joshi, José Cambronero Sanchez, Sumit Gulwani, Vu Le, Gust Verbruggen, and Ivan Radiček. 2023 · 2023
Cited alongside, same era.
On the Impact of Language Selection for Training and Evaluating Programming Language Models. In 2023 IEEE 23rd International Working Conference on Source Code Analysis and Manipulation (SCAM) . IEEE Computer Society, Los Alamitos, CA, USA, 271–276
J. Katzy, M. Izadi, and A. Deursen. 2023 · 2023
Cited alongside, same era.
Max Schäfer, Sarah Nadi, Aryaz Eghbali, and Frank Tip. 2024 · 2023
Closest in time.
AI Assistant for Software Developers | Tabnine
TabNine. 2023 · 2023
Closest in time.
Self-Instruct: Aligning Language Model with Self Generated Instructions. In Annual Meeting of the Association of Computation Linguistics (ACL)
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 · 2023
Closest in time.
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Evaluations on HumanEval-X. In ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . Association for Computing Machinery, 5673–5684
Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023 · 2023
Closest in time.
Big Code Models Leaderboard
Loubna Ben Allal. 2024 · 2024
Closest in time.
StudentEval: A Benchmark of Student-Written Prompts for Large Language Models of Code. In Findings of the Association for Computational Linguistics
Hannah McLean Babe, Sydney Nguyen, Yangtian Zi, Arjun Guha, Molly Q. Feldman, and Carolyn Jane Anderson. 2024 · 2024
Closest in time.
Learning Transfers over Several Programming Languages
Razan Baltaji, Saurabh Pujar, Louis Mandel, Martin Hirzel, Luca Buratti, and Lav Varshney. 2024 · 2024
Closest in time.
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. In International Conference on Machine Learning (ICML)
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu. 2024 · 2024
Closest in time.
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions. In International Workshop on Large Language Models for Code (LLM4Code)
Federico Cassano, Luisa Li, Akul Sethi, Noah Shinn, Abby Brennan-Jones, Anton Lozhkov, Carolyn Jane Anderson, and Arjun Guha. 2024 · 2024
Closest in time.
DeepSeek-Coder: When the Large Language Model Meets Programming – The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024 · 2024
Closest in time.
StarCoder 2 and The Stack v2: The Next Generation
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo, Evgenii Zheltonozhskii, Nii Osae Osae Dade, Wenhao Yu, Lucas Krauß, Naman Jain, Yixuan Su, Xuanli He, Manan Dey, Edoardo Abati, Yekun Chai, Niklas Muennighoff, Xiangru Tang, Muhtasham Oblokulov, Christopher Akiki, Marc Marone, Chenghao Mou, Mayank Mishra, Alex Gu, Binyuan Hui, Tri Dao, Armel Zebaze, Olivier Dehaene, Nicolas Patry, Canwen Xu, Julian McAuley, Han Hu, Torsten Scholak, Sebastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, Mostofa Patwary, Nima Tajbakhsh, Yacine Jernite, Carlos Muñoz Ferrandis, Lingming Zhang, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2024 · 2024
Closest in time.
WizardCoder: Empowering Code Large Language Models with Evol-Instruct. In The Twelfth International Conference on Learning Representations
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2024 · 2024
Closest in time.
OctoPack: Instruction Tuning Code Large Language Models. In International Conference on Learning Representations (ICLR)
Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, and Shayne Longpre. 2024 · 2024
Closest in time.
Using an LLM to Help With Code Understanding. In IEEE/ACM 46th International Conference on Software Engineering (ICSE) . Association for Computing Machinery, 1–13
Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024 · 2024
Closest in time.
Lost in Translation: A Study of Bugs Introduced by Large Language Models While Translating Code. In IEEE/ACM International Conference on Software Engineering (ICSE) (ICSE ’24) . Association for Computing Machinery, New York, NY, USA, 1–13
Rangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar, Lambert Pouguem Wassi, Michele Merler, Boris Sobolev, Raju Pavuluri, Saurabh Sinha, and Reyhaneh Jabbarvand. 2024 · 2024
Closest in time.
Magicoder: Empowering Code Generation with OSS-Instruct. In International Conference on Machine Learning (ICML)
Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. 2024 · 2024
Closest in time.
Fuzz4All: Universal Fuzzing with Large Language Models. In IEEE/ACM International Conference on Software Engineering (ICSE) . Association for Computing Machinery, New York, NY, USA, 1–13
Chunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel, and Lingming Zhang. 2024 · 2024
Closest in time.
Data Augmentation for Code Translation with Comparable Corpora and Multiple References. In Findings of EMNLP
Yiqing Xie, Atharva Naik, Daniel Fried, and Carolyn Rose. 2024 · 2024
Closest in time.