Fetching the paper…
Reading the bibliography…
In this work we systematically review the recent advancements in software engineering with language models, covering 70+ models, 40+ evaluation tasks, 180+ datasets, and 900 related works.
Predicting variable types in dynamically typed programming languages
Abhinav Jangda and Gaurav Anand · 1901
Earlier work this paper cites.
Uniparser: A unified log parser for heterogeneous log data
Yudong Liu, Xu Zhang, Shilin He, Hongyu Zhang, Liqun Li, Yu Kang, Yong Xu, Minghua Ma, Qingwei Lin, Yingnong Dang, Saravan Rajmohan, and Dongmei Zhang · 1901
Earlier work this paper cites.
A comprehensive exploration on wikisql with table-aware word contextualization
Wonseok Hwang, Jinyeung Yim, Seunghyun Park, and Minjoon Seo · 1902
Earlier work this paper cites.
Sketch2code: Generating a website from a paper mockup
Alex Robinson · 1905
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Neural code search evaluation dataset
Hongyu Li, Seohyun Kim, and Satish Chandra · 1908
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 1909
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer · 1911
Earlier work this paper cites.
Dltpy: Deep learning type inference of python function signatures using natural language context
Casper Boone, Niels de Bruin, Arjan Langerak, and Fabian Stelmach · 1912
Earlier work this paper cites.
Cosql: A conversational text-to-sql challenge towards cross-domain natural language interfaces to databases
Tao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander R. Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter S. Lasecki, and Dragomir R. Radev · 1979
Earlier work this paper cites.
The ATIS spoken language systems pilot corpus
Charles T. Hemphill, John J. Godfrey, and George R. Doddington · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Expanding the scope of the ATIS task: The ATIS-3 corpus
Deborah A. Dahl, Madeleine Bates, Michael Brown, William M. Fisher, Kate Hunicke-Smith, David S. Pallett, Christine Pao, Alexander I. Rudnicky, and Elizabeth Shriberg · 1994
Earlier work this paper cites.
Learning to parse database queries using inductive logic programming
John M. Zelle and Raymond J. Mooney · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Automated construction of database interfaces: Intergrating statistical and relational learning for semantic parsing
Lappoon R. Tang and Raymond J. Mooney · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Watermarking, tamper-proofing, and obfuscation - tools for software protection
Christian S. Collberg and Clark Thomborson · 2002
Earlier work this paper cites.
Bertrand-dr: Improving text-to-sql using a discriminative re-ranker
Amol Kelkar, Rohan Relan, Vaishali Bhardwaj, Saurabh Vaichal, and Peter Relan · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Transˆ3: A transformer-based framework for unifying code summarization and code search
Wenhua Wang, Yuqun Zhang, Zhengran Zeng, and Guandong Xu · 2003
Earlier work this paper cites.
LLVM: A compilation framework for lifelong program analysis & transformation
Chris Lattner and Vikram S. Adve · 2004
Earlier work this paper cites.
Bao: Learning to steer query optimizers
Ryan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul, Mohammad Alizadeh, and Tim Kraska · 2004
Earlier work this paper cites.
Opttyper: Probabilistic type inference by optimising logical and natural constraints
Irene Vlassi Pandi, Earl T. Barr, Andrew D. Gordon, and Charles Sutton · 2004
Earlier work this paper cites.
Natural language processing (NLP) for requirements engineering: A systematic mapping study
Liping Zhao, Waad Alhoshan, Alessio Ferrari, Keletso J. Letsholo, Muideen A. Ajagbe, Erol-Valeriu Chioasca, and Riza Theresa Batista-Navarro · 2004
Earlier work this paper cites.
Scalable inference and training of context-rich syntactic translation models
Michel Galley, Jonathan Graehl, Kevin Knight, Daniel Marcu, Steve DeNeefe, Wei Wang, and Ignacio Thayer · 2006
Earlier work this paper cites.
A survey on the evaluation of clone detection performance and benchmarking
Jeffrey Svajlenko and Chanchal K. Roy · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2006
Earlier work this paper cites.
Hierarchical phrase-based translation
David Chiang · 2007
Earlier work this paper cites.
DECKARD: scalable and accurate tree-based detection of code clones
Lingxiao Jiang, Ghassan Misherghi, Zhendong Su, and Stéphane Glondu · 2007
Earlier work this paper cites.
Loghub: A large collection of system log datasets towards automated log analytics
Shilin He, Jieming Zhu, Pinjia He, and Michael R. Lyu · 2008
Earlier work this paper cites.
Neural code search revisited: Enhancing code snippet retrieval through natural language intent
Geert Heyman and Tom Van Cutsem · 2008
Earlier work this paper cites.
Static uml model generator from analysis of requirements (sugar)
Deeptimahanti Deva Kumar and Ratna Sanyal · 2008
Earlier work this paper cites.
Hybrid ranking network for text-to-sql
Qin Lyu, Kaushik Chakrabarti, Shobhit Hathi, Souvik Kundu, Jianwen Zhang, and Zheng Chen · 2008
Earlier work this paper cites.
Learning from examples to improve code completion systems
Marcel Bruch, Martin Monperrus, and Mira Mezini · 2009
Earlier work this paper cites.
An automated tool for generating UML models from natural language requirements
Deva Kumar Deeptimahanti and Muhammad Ali Babar · 2009
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma · 2009
Earlier work this paper cites.
Class diagram extraction from textual requirements using natural language processing (nlp) techniques
Mohd Ibrahim and Rodina Ahmad · 2010
Earlier work this paper cites.
Evading virus detection using code obfuscation
Khurram Murad, Syed Noor-ul-Hassan Shirazi, Yousaf Bin Zikria, and Nassar Ikram · 2010
Earlier work this paper cites.
The significance of user-defined identifiers in java source code authorship identification
Georgia Frantzeskou, Stephen G. MacDonell, Efstathios Stamatatos, Stelios Georgiou, and Stefanos Gritzalis · 2011
Earlier work this paper cites.
Evosuite: automatic test suite generation for object-oriented software
Gordon Fraser and Andrea Arcuri · 2011
Earlier work this paper cites.
From user requirements to UML class diagram
Hatem Herchi and Wahiba Ben Abdessalem · 2012
Earlier work this paper cites.
On the naturalness of software
Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, and Premkumar T. Devanbu · 2012
Earlier work this paper cites.
Mining source code repositories at massive scale using language modeling
Miltiadis Allamanis and Charles Sutton · 2013
Earlier work this paper cites.
Reducing human effort and improving quality in peer code reviews using automatic static analysis and reviewer recommendation
Vipin Balachandran · 2013
Earlier work this paper cites.
Deep neural networks for source code author identification
Upul Bandara and Gamini Wijayarathna · 2013
Earlier work this paper cites.
TRAM: A tool for transforming textual requirements into analysis models
Keletso Letsholo, Liping Zhao, and Erol-Valeriu Chioasca · 2013
Earlier work this paper cites.
A machine learning framework for programming by example
Aditya Krishna Menon, Omer Tamuz, Sumit Gulwani, Butler W. Lampson, and Adam Kalai · 2013
Earlier work this paper cites.
Lexical statistical machine translation for language migration
Anh Tuan Nguyen, Tung Thanh Nguyen, and Tien N. Nguyen · 2013
Earlier work this paper cites.
Autocomment: Mining question and answer sites for automatic comment generation
Edmund Wong, Jinqiu Yang, and Lin Tan · 2013
Earlier work this paper cites.
Mining idioms from source code
Miltiadis Allamanis and Charles Sutton · 2014
Earlier work this paper cites.
Learning natural coding conventions
Miltiadis Allamanis, Earl T. Barr, Christian Bird, and Charles Sutton · 2014
Earlier work this paper cites.
On automatically generating commit messages via summarization of source code changes
Luis Fernando Cortes-Coy, Mario Linares Vásquez, Jairo Aponte, and Denys Poshyvanyk · 2014
Earlier work this paper cites.
Measuring and improving the completeness of natural language requirements
Alessio Ferrari, Felice Dell’Orletta, Giorgio Oronzo Spagnolo, and Stefania Gnesi · 2014
Earlier work this paper cites.
A large-scale evaluation of automated unit test generation using evosuite
Gordon Fraser and Andrea Arcuri · 2014
Earlier work this paper cites.
The major mutation framework: efficient and scalable mutation analysis for java
René Just · 2014
Earlier work this paper cites.
Defects4j: a database of existing faults to enable controlled testing studies for java programs
René Just, Darioush Jalali, and Michael D. Ernst · 2014
Earlier work this paper cites.
Phrase-based statistical translation of programming languages
Svetoslav Karaivanov, Veselin Raychev, and Martin T. Vechev · 2014
Earlier work this paper cites.
Constructing an interactive natural language interface for relational databases
Fei Li and H. V. Jagadish · 2014
Earlier work this paper cites.
Code completion with statistical language models
Veselin Raychev, Martin T. Vechev, and Eran Yahav · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Towards a big data curated benchmark of inter-project code clones
Jeffrey Svajlenko, Judith F. Islam, Iman Keivanloo, Chanchal Kumar Roy, and Mohammad Mamun Mia · 2014
Earlier work this paper cites.
On the localness of software
Zhaopeng Tu, Zhendong Su, and Premkumar T. Devanbu · 2014
Earlier work this paper cites.
Suggesting accurate method and class names
Miltiadis Allamanis, Earl T. Barr, Christian Bird, and Charles Sutton · 2015
Earlier work this paper cites.
Automated checking of conformance to requirements templates using natural language processing
Chetan Arora, Mehrdad Sabetzadeh, Lionel C. Briand, and Frank Zimmer · 2015
Earlier work this paper cites.
Change impact analysis for natural language requirements: An NLP approach
Chetan Arora, Mehrdad Sabetzadeh, Arda Goknil, Lionel C. Briand, and Frank Zimmer · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Helping developers help themselves: Automatic decomposition of code review changesets
Mike Barnett, Christian Bird, João Brunet, and Shuvendu K. Lahiri · 2015
Earlier work this paper cites.
Program-adaptive mutational fuzzing
Sang Kil Cha, Maverick Woo, and David Brumley · 2015
Earlier work this paper cites.
The manybugs and introclass benchmarks for automated repair of C programs
Claire Le Goues, Neal J. Holtschulte, Edward K. Smith, Yuriy Brun, Premkumar T. Devanbu, Stephanie Forrest, and Westley Weimer · 2015
Earlier work this paper cites.
Query expansion via wordnet for effective code search
Meili Lu, Xiaobing Sun, Shaowei Wang, David Lo, and Yucong Duan · 2015
Earlier work this paper cites.
Codehow: Effective code search based on API understanding and extended boolean model (E)
Fei Lv, Hongyu Zhang, Jian-Guang Lou, Shaowei Wang, Dongmei Zhang, and Jianjun Zhao · 2015
Earlier work this paper cites.
Divide-and-conquer approach for multi-phase statistical migration for source code (T)
Anh Tuan Nguyen, Tung Thanh Nguyen, and Tien N. Nguyen · 2015
Earlier work this paper cites.
Typedevil: Dynamic type inconsistency analysis for javascript
Michael Pradel, Parker Schuh, and Koushik Sen · 2015
Earlier work this paper cites.
Predicting program properties from "big code"
Veselin Raychev, Martin T. Vechev, and Andreas Krause · 2015
Earlier work this paper cites.
Tricorder: Building a program analysis ecosystem
Caitlin Sadowski, Jeffrey van Gogh, Ciera Jaspan, Emma Söderberg, and Collin Winter · 2015
Earlier work this paper cites.
Automated unit test generation for evolving software
Sina Shamshiri · 2015
Earlier work this paper cites.
Changescribe: A tool for automatically generating commit messages
Mario Linares Vásquez, Luis Fernando Cortes-Coy, Jairo Aponte, and Denys Poshyvanyk · 2015
Earlier work this paper cites.
Toward deep learning software repositories
Martin White, Christopher Vendome, Mario Linares Vásquez, and Denys Poshyvanyk · 2015
Earlier work this paper cites.
Clocom: Mining existing source code for automatic comment generation
Edmund Wong, Taiyue Liu, and Lin Tan · 2015
Earlier work this paper cites.
Extracting domain models from natural-language requirements: approach and industrial evaluation
Chetan Arora, Mehrdad Sabetzadeh, Lionel C. Briand, and Frank Zimmer · 2016
Earlier work this paper cites.
Automated extraction and clustering of requirements glossary terms
Chetan Arora, Mehrdad Sabetzadeh, Lionel C. Briand, and Frank Zimmer · 2016
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Extracting features from online software reviews to aid requirements reuse
Noor Hasrina Bakar, Zarinah Mohd Kasirun, Norsaremah Salleh, and Hamid Abdullah Jalab · 2016
Earlier work this paper cites.
Statistical deobfuscation of android applications
Benjamin Bichsel, Veselin Raychev, Petar Tsankov, and Martin T. Vechev · 2016
Earlier work this paper cites.
PHOG: probabilistic model for code
Pavol Bielik, Veselin Raychev, and Martin T. Vechev · 2016
Earlier work this paper cites.
PIT: a practical mutation testing tool for java (demo)
Henry Coles, Thomas Laurent, Christopher Henard, Mike Papadakis, and Anthony Ventresque · 2016
Earlier work this paper cites.
LAVA: large-scale automated vulnerability addition
Brendan Dolan-Gavitt, Patrick Hulin, Engin Kirda, Tim Leek, Andrea Mambretti, William K. Robertson, Frederick Ulrich, and Ryan Whelan · 2016
Earlier work this paper cites.
Spell: Streaming parsing of system event logs
Min Du and Feifei Li · 2016
Earlier work this paper cites.
ARSENAL: automatic requirements specification extraction from natural language
Shalini Ghosh, Daniel Elenius, Wenchao Li, Patrick Lincoln, Natarajan Shankar, and Wilfried Steiner · 2016
Earlier work this paper cites.
Discovering bug patterns in javascript
Quinn Hanam, Fernando Santos De Mattos Brito, and Ali Mesbah · 2016
Earlier work this paper cites.
Summarizing source code using a neural attention model
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer · 2016
Earlier work this paper cites.
Relationship-aware code search for javascript frameworks
Xuan Li, Zerui Wang, Qianxiang Wang, Shoumeng Yan, Tao Xie, and Hong Mei · 2016
Earlier work this paper cites.
Latent predictor networks for code generation
Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomás Kociský, Fumin Wang, and Andrew W. Senior · 2016
Earlier work this paper cites.
Automatic patch generation by learning correct code
Fan Long and Martin C. Rinard · 2016
Earlier work this paper cites.
Convolutional neural networks over tree structures for programming language processing
Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin · 2016
Earlier work this paper cites.
Mapping API elements for code migration with vector representations
Trong Duc Nguyen, Anh Tuan Nguyen, and Tien N. Nguyen · 2016
Earlier work this paper cites.
Evilcoder: automated bug insertion
Jannik Pewny and Thorsten Holz · 2016
Earlier work this paper cites.
sk_p: a neural program corrector for moocs
Yewen Pu, Karthik Narasimhan, Armando Solar-Lezama, and Regina Barzilay · 2016
Earlier work this paper cites.
On the "naturalness" of buggy code
Baishakhi Ray, Vincent J. Hellendoorn, Saheel Godhane, Zhaopeng Tu, Alberto Bacchelli, and Premkumar T. Devanbu · 2016
Earlier work this paper cites.
Probabilistic model for code with decision trees
Veselin Raychev, Pavol Bielik, and Martin T. Vechev · 2016
Earlier work this paper cites.
Learning programs from noisy data
Veselin Raychev, Pavol Bielik, Martin T. Vechev, and Andreas Krause · 2016
Earlier work this paper cites.
Automated extraction of conceptual models from user stories via NLP
Marcel Robeer, Garm Lucassen, Jan Martijn E. M. van der Werf, Fabiano Dalpiaz, and Sjaak Brinkkemper · 2016
Earlier work this paper cites.
Sourcerercc: scaling code clone detection to big-code
Hitesh Sajnani, Vaibhav Saini, Jeffrey Svajlenko, Chanchal K. Roy, and Cristina V. Lopes · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Bugram: bug detection with n-gram language models
Song Wang, Devin Chollak, Dana Movshovitz-Attias, and Lin Tan · 2016
Earlier work this paper cites.
Automatically learning semantic features for defect prediction
Song Wang, Taiyue Liu, and Lin Tan · 2016
Earlier work this paper cites.
Deep learning code fragments for code clone detection
Martin White, Michele Tufano, Christopher Vendome, and Denys Poshyvanyk · 2016
Earlier work this paper cites.
Python probabilistic type inference with natural language support
Zhaogui Xu, Xiangyu Zhang, Lin Chen, Kexin Pei, and Baowen Xu · 2016
Earlier work this paper cites.
Integrating semantic nlp and logic reasoning into a unified system for fully-automated code checking
Jiansong Zhang and Nora M. El-Gohary · 2016
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron C. Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
Deepcoder: Learning to write programs
Matej Balog, Alexander L. Gaunt, Marc Brockschmidt, Sebastian Nowozin, and Daniel Tarlow · 2017
Earlier work this paper cites.
Directed greybox fuzzing
Marcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, and Abhik Roychoudhury · 2017
Earlier work this paper cites.
Coverage-based greybox fuzzing as markov chain
Marcel Böhme, Van-Thuan Pham, and Abhik Roychoudhury · 2017
Earlier work this paper cites.
The care and feeding of wild-caught mutants
David Bingham Brown, Michael Vaughn, Ben Liblit, and Thomas W. Reps · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
End-to-end deep learning of optimization heuristics
Chris Cummins, Pavlos Petoumenos, Zheng Wang, and Hugh Leather · 2017
Earlier work this paper cites.
Robustfill: Neural program learning under noisy I/O
Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, and Pushmeet Kohli · 2017
Earlier work this paper cites.
Deeplog: Anomaly detection and diagnosis from system logs through deep learning
Min Du, Feifei Li, Guineng Zheng, and Vivek Srikumar · 2017
Earlier work this paper cites.
Deepam: Migrate apis with multi-modal sequence to sequence learning
Xiaodong Gu, Hongyu Zhang, Dongmei Zhang, and Sunghun Kim · 2017
Earlier work this paper cites.
Deepfix: Fixing common C language errors by deep learning
Rahul Gupta, Soham Pal, Aditya Kanade, and Shirish K. Shevade · 2017
Earlier work this paper cites.
A little bird told me: Mining tweets for requirements and software evolution
Emitza Guzman, Mohamed Ibrahim, and Martin Glinz · 2017
Earlier work this paper cites.
Drain: An online log parsing approach with fixed depth tree
Pinjia He, Jieming Zhu, Zibin Zheng, and Michael R. Lyu · 2017
Earlier work this paper cites.
Are deep neural networks the best choice for modeling source code?
Vincent J. Hellendoorn and Premkumar T. Devanbu · 2017
Earlier work this paper cites.
Learning a neural semantic parser from user feedback
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Towards automatic generation of short summaries of commits
Siyuan Jiang and Collin McMillan · 2017
Earlier work this paper cites.
Automatically generating commit messages from diffs using neural machine translation
Siyuan Jiang, Ameer Armaly, and Collin McMillan · 2017
Earlier work this paper cites.
SAFE: A simple approach for feature extraction from app descriptions and app reviews
Timo Johann, Christoph Stanik, Alireza M. Alizadeh B., and Walid Maalej · 2017
Earlier work this paper cites.
Quixbugs: a multi-lingual program repair benchmark set based on the quixey challenge
Derrick Lin, James Koppel, Angela Chen, and Armando Solar-Lezama · 2017
Earlier work this paper cites.
Automatic classification of non-functional requirements from augmented app user reviews
Mengmeng Lu and Peng Liang · 2017
Earlier work this paper cites.
Learning to infer API mappings from API documents
Yangyang Lu, Ge Li, Zelong Zhao, Linfeng Wen, and Zhi Jin · 2017
Earlier work this paper cites.
Automated test case generation as a many-objective optimisation problem with dynamic selection of the targets
Annibale Panichella, Fitsum Meshesha Kifetew, and Paolo Tonella · 2017
Earlier work this paper cites.
Neuro-symbolic program synthesis
Emilio Parisotto, Abdel-rahman Mohamed, Rishabh Singh, Lihong Li, Dengyong Zhou, and Pushmeet Kohli · 2017
Earlier work this paper cites.
Statistical migration of API usages
Hung Dang Phan, Anh Tuan Nguyen, Trong Duc Nguyen, and Tien N. Nguyen · 2017
Earlier work this paper cites.
Abstract syntax networks for code generation and semantic parsing
Maxim Rabinovich, Mitchell Stern, and Dan Klein · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Codeflaws: a programming competition benchmark for evaluating automated program repair tools
Shin Hwei Tan, Jooyong Yi, Yulis, Sergey Mechtaev, and Abhik Roychoudhury · 2017
Earlier work this paper cites.
Recovering clear, natural identifiers from obfuscated JS names
Bogdan Vasilescu, Casey Casalnuovo, and Premkumar T. Devanbu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Supervised deep features for software functional clone detection by exploiting lexical and syntactical information in source code
Huihui Wei and Ming Li · 2017
Earlier work this paper cites.
Sqlnet: Generating structured queries from natural language without reinforcement learning
Xiaojun Xu, Chang Liu, and Dawn Song · 2017
Earlier work this paper cites.
Sqlizer: query synthesis from natural language
Navid Yaghmazadeh, Yuepeng Wang, Isil Dillig, and Thomas Dillig · 2017
Earlier work this paper cites.
A syntactic neural model for general-purpose code generation
Pengcheng Yin and Graham Neubig · 2017
Earlier work this paper cites.
Recommending apis for API related questions in stack overflow
Jingxuan Zhang, He Jiang, Zhilei Ren, and Xin Chen · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Victor Zhong, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Large-scale and language-oblivious code authorship identification
Mohammed Abuhamad, Tamer AbuHmed, Aziz Mohaisen, and DaeHun Nyang · 2018
Earlier work this paper cites.
Learning to represent programs with graphs
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi · 2018
Earlier work this paper cites.
A general path-based representation for predicting program properties
Uri Alon, Meital Zilberstein, Omer Levy, and Eran Yahav · 2018
Earlier work this paper cites.
pix2code: Generating code from a graphical user interface screenshot
Tony Beltramelli · 2018
Earlier work this paper cites.
Neural code comprehension: A learnable representation of code semantics
Tal Ben-Nun, Alice Shoshana Jakobovits, and Torsten Hoefler · 2018
Earlier work this paper cites.
Neuro-symbolic program corrector for introductory programming assignments
Sahil Bhatia, Pushmeet Kohli, and Rishabh Singh · 2018
Earlier work this paper cites.
Nar-miner: discovering negative association rules from code for bug detection
Pan Bian, Bin Liang, Wenchang Shi, Jianjun Huang, and Yan Cai · 2018
Earlier work this paper cites.
Leveraging grammar and reinforcement learning for neural program synthesis
Rudy Bunel, Matthew J. Hausknecht, Jacob Devlin, Rishabh Singh, and Pushmeet Kohli · 2018
Earlier work this paper cites.
Angora: Efficient fuzzing by principled search
Peng Chen and Hao Chen · 2018
Earlier work this paper cites.
Tree-to-tree neural networks for program translation
Xinyun Chen, Chang Liu, and Dawn Song · 2018
Earlier work this paper cites.
Coarse-to-fine decoding for neural semantic parsing
Li Dong and Mirella Lapata · 2018
Earlier work this paper cites.
Program language translation using a grammar-driven tree-to-tree model
Mehdi Drissi, Olivia Watkins, Aditya Khant, Vivaswat Ojha, Pedro Sandoval Segura, Rakia Segev, Eric Weiner, and Robert Keller · 2018
Earlier work this paper cites.
Automatic transformation of user stories into UML use case diagrams using NLP techniques
Meryem Elallaoui, Khalid Nafil, and Raja Touahni · 2018
Earlier work this paper cites.
Program synthesis using conflict-driven learning
Yu Feng, Ruben Martins, Osbert Bastani, and Isil Dillig · 2018
Earlier work this paper cites.
Improving text-to-sql evaluation methodology
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir R. Radev · 2018
Earlier work this paper cites.
Automatic software repair: a survey
Luca Gazzola, Daniela Micucci, and Leonardo Mariani · 2018
Earlier work this paper cites.
Deep code search
Xiaodong Gu, Hongyu Zhang, and Sunghun Kim · 2018
Earlier work this paper cites.
Intelligent code reviews using deep learning, 2018
Anshul Gupta and Neel Sundaresan · 2018
Earlier work this paper cites.
Maxsmt-based type inference for python 3
Mostafa Hassan, Caterina Urban, Marco Eilers, and Peter Müller · 2018
Earlier work this paper cites.
Retrieval-based neural code generation
Shirley Anugrah Hayati, Raphaël Olivier, Pravalika Avvaru, Pengcheng Yin, Anthony Tomasic, and Graham Neubig · 2018
Earlier work this paper cites.
Debin: Predicting debug information in stripped binaries
Jingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev, and Martin T. Vechev · 2018
Earlier work this paper cites.
Deep learning type inference
Vincent J. Hellendoorn, Christian Bird, Earl T. Barr, and Miltiadis Allamanis · 2018
Earlier work this paper cites.
Deep code comment generation
Xing Hu, Ge Li, Xin Xia, David Lo, and Zhi Jin · 2018
Earlier work this paper cites.
Summarizing source code with transferred API knowledge
Xing Hu, Ge Li, Xin Xia, David Lo, Shuai Lu, and Zhi Jin · 2018
Earlier work this paper cites.
API method recommendation without worrying about the task-api knowledge gap
Qiao Huang, Xin Xia, Zhenchang Xing, David Lo, and Xinyu Wang · 2018
Earlier work this paper cites.
Mapping language to code in programmatic context
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Neural-guided deductive search for real-time program synthesis from examples
Ashwin Kalyan, Abhishek Mohta, Oleksandr Polozov, Dhruv Batra, Prateek Jain, and Sumit Gulwani · 2018
Earlier work this paper cites.
Facoy: a code-to-code search engine
Kisub Kim, Dongsun Kim, Tegawendé F. Bissyandé, Eunjong Choi, Li Li, Jacques Klein, and Yves Le Traon · 2018
Earlier work this paper cites.
Accelerating search-based program synthesis using learned probabilistic models
Woosuk Lee, Kihong Heo, Rajeev Alur, and Mayur Naik · 2018
Earlier work this paper cites.
Fairfuzz: a targeted mutation strategy for increasing greybox fuzz testing coverage
Caroline Lemieux and Koushik Sen · 2018
Earlier work this paper cites.
Code completion with neural attention and pointer networks
Jian Li, Yue Wang, Michael R. Lyu, and Irwin King · 2018
Earlier work this paper cites.
Vuldeepecker: A deep learning-based system for vulnerability detection
Zhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin, Sujuan Wang, Zhijun Deng, and Yuyi Zhong · 2018
Earlier work this paper cites.
Cross-project transfer representation learning for vulnerable function discovery
Guanjun Lin, Jun Zhang, Wei Luo, Lei Pan, Yang Xiang, Olivier Y. de Vel, and Paul Montague · 2018
Earlier work this paper cites.
Nl2bash: A corpus and semantic parser for natural language interface to the linux operating system
Xi Victoria Lin, Chenglong Wang, Luke Zettlemoyer, and Michael D. Ernst · 2018
Earlier work this paper cites.
Neural-machine-translation-based commit message generation: how far are we?
Zhongxin Liu, Xin Xia, Ahmed E. Hassan, David Lo, Zhenchang Xing, and Xinyu Wang · 2018
Earlier work this paper cites.
Content aware source code change description generation
Pablo Loyola, Edison Marrese-Taylor, Jorge A. Balazs, Yutaka Matsuo, and Fumiko Satoh · 2018
Earlier work this paper cites.
Detecting anomaly in big data system logs using convolutional neural network
Siyang Lu, Xiang Wei, Yandong Li, and Liqiang Wang · 2018
Earlier work this paper cites.
Industrial requirements classification for redundancy and inconsistency detection in SEMIOS
Manel Mezghani, Juyeon Kang, and Florence Sèdes · 2018
Earlier work this paper cites.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Earlier work this paper cites.
Automatic software repair: A bibliography
Martin Monperrus · 2018
Earlier work this paper cites.
Building language models for text with named entities
Md. Rizwan Parvez, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Deepbugs: a learning approach to name-based bug detection
Michael Pradel and Koushik Sen · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Bug synthesis: challenging bug-finding tools with deep faults
Subhajit Roy, Awanish Pandey, Brendan Dolan-Gavitt, and Yu Hu · 2018
Earlier work this paper cites.
Automated vulnerability detection in source code using deep representation learning
Rebecca L. Russell, Louis Y. Kim, Lei H. Hamilton, Tomo Lazovich, Jacob Harer, Onur Ozdemir, Paul M. Ellingwood, and Marc W. McConley · 2018
Earlier work this paper cites.
Retrieval on source code: a neural code search
Saksham Sachdev, Hongyu Li, Sifei Luan, Seohyun Kim, Koushik Sen, and Satish Chandra · 2018
Earlier work this paper cites.
Bugs.jar: a large-scale, diverse dataset of real-world java bugs
Ripon K. Saha, Yingjun Lyu, Wing Lam, Hiroaki Yoshida, and Mukul R. Prasad · 2018
Earlier work this paper cites.
Oreo: detection of clones in the twilight zone
Vaibhav Saini, Farima Farmahinifarahani, Yadong Lu, Pierre Baldi, and Cristina V. Lopes · 2018
Earlier work this paper cites.
Test generation for higher-order functions in dynamic languages
Marija Selakovic, Michael Pradel, Rezwana Karim, and Frank Tip · 2018
Earlier work this paper cites.
Learning to map context-dependent sentences to executable formal queries
Alane Suhr, Srinivasan Iyer, and Yoav Artzi · 2018
Earlier work this paper cites.
Improving automatic source code summarization via deep reinforcement learning
Yao Wan, Zhou Zhao, Min Yang, Guandong Xu, Haochao Ying, Jian Wu, and Philip S. Yu · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2018
Earlier work this paper cites.
Automated generation of constraints from use case specifications to support system testing
Chunhui Wang, Fabrizio Pastore, and Lionel C. Briand · 2018
Earlier work this paper cites.
Ccaligner: a token based large-gap clone detector
Pengcheng Wang, Jeffrey Svajlenko, Yanzhao Wu, Yun Xu, and Chanchal K. Roy · 2018
Earlier work this paper cites.
Machine learning in compiler optimisation
Zheng Wang and Michael F. P. O’Boyle · 2018
Earlier work this paper cites.
Staqc: A systematically mined question-code dataset from stack overflow
Ziyu Yao, Daniel S. Weld, Wei-Peng Chen, and Huan Sun · 2018
Earlier work this paper cites.
Learning to mine aligned code and natural language pairs from stack overflow
Pengcheng Yin, Bowen Deng, Edgar Chen, Bogdan Vasilescu, and Graham Neubig · 2018
Earlier work this paper cites.
Syntaxsqlnet: Syntax tree networks for complex and cross-domain text-to-sql task
Tao Yu, Michihiro Yasunaga, Kai Yang, Rui Zhang, Dongxu Wang, Zifan Li, and Dragomir R. Radev · 2018
Earlier work this paper cites.
Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir R. Radev · 2018
Earlier work this paper cites.
Deepsim: deep learning code functional similarity
Gang Zhao and Jeff Huang · 2018
Earlier work this paper cites.
Juice: A large scale distantly supervised dataset for open domain context-based code generation
Rajas Agashe, Srinivasan Iyer, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
code2seq: Generating sequences from structured representations of code
Uri Alon, Shaked Brody, Omer Levy, and Eran Yahav · 2019
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi · 2019
Earlier work this paper cites.
Automatic html code generation from mock-up images using machine learning techniques
Batuhan Aşıroğlu, Büşta Rümeysa Mete, Eyyüp Yıldız, Yağız Nalçakan, Alper Sezen, Mustafa Dağtekin, and Tolga Ensari · 2019
Earlier work this paper cites.
Autopandas: neural-backed generators for program synthesis
Rohan Bavishi, Caroline Lemieux, Roy Fox, Koushik Sen, and Ion Stoica · 2019
Earlier work this paper cites.
Representing schema structure with graph neural networks for text-to-sql parsing
Ben Bogin, Jonathan Berant, and Matt Gardner · 2019
Earlier work this paper cites.
SAR: learning cross-language API mappings with little knowledge
Nghi D. Q. Bui, Yijun Yu, and Lingxiao Jiang · 2019
Earlier work this paper cites.
When deep learning met code search
José Cambronero, Hongyu Li, Seohyun Kim, Koushik Sen, and Satish Chandra · 2019
Earlier work this paper cites.
Mining likely analogical apis across third-party libraries via large-scale unsupervised API semantics embedding
Chunyang Chen, Zhenchang Xing, Yang Liu, and Kent Ong Long Xiong · 2019
Earlier work this paper cites.
Sequencer: Sequence-to-sequence learning for end-to-end program repair
Zimin Chen, Steve Kommrusch, Michele Tufano, Louis-Noël Pouchet, Denys Poshyvanyk, and Martin Monperrus · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2019
Earlier work this paper cites.
Structured neural summarization
Patrick Fernandes, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Towards complex text-to-sql in cross-domain database with intermediate representation
Jiaqi Guo, Zecheng Zhan, Yan Gao, Yan Xiao, Jian-Guang Lou, Ting Liu, and Dongmei Zhang · 2019
Earlier work this paper cites.
Bugsjs: a benchmark of javascript bugs
Péter Gyimesi, Béla Vancsics, Andrea Stocco, Davood Mazinanian, Árpád Beszédes, Rudolf Ferenc, and Ali Mesbah · 2019
Earlier work this paper cites.
Re-factoring based program repair applied to programming assignments
Yang Hu, Umair Z. Ahmed, Sergey Mechtaev, Ben Leong, and Abhik Roychoudhury · 2019
Earlier work this paper cites.
Machine learning based recommendation of method names: How far are we
Lin Jiang, Hui Liu, and He Jiang · 2019
Earlier work this paper cites.
Code authorship attribution: Methods and challenges
Vaibhavi Kalgutkar, Ratinder Kaur, Hugo Gonzalez, Natalia Stakhanova, and Alina Matyukhina · 2019
Earlier work this paper cites.
Automated customized bug-benchmark generation
Vineeth Kashyap, Jason Ruchti, Lucja Kot, Emma Turetsky, Rebecca Swords, Shih An Pan, Julien Henry, David Melski, and Eric M. Schulte · 2019
Earlier work this paper cites.
Spoc: Search-based pseudocode to code
Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, and Percy Liang · 2019
Earlier work this paper cites.
DIRE: A neural approach to decompiled identifier naming
Jeremy Lacomis, Pengcheng Yin, Edward J. Schwartz, Miltiadis Allamanis, Claire Le Goues, Graham Neubig, and Bogdan Vasilescu · 2019
Earlier work this paper cites.
A neural model for generating natural language summaries of program subroutines
Alexander LeClair, Siyuan Jiang, and Collin McMillan · 2019
Earlier work this paper cites.
Unsupervised pivot translation for distant languages
Yichong Leng, Xu Tan, Tao Qin, Xiang-Yang Li, and Tie-Yan Liu · 2019
Earlier work this paper cites.
Deepreview: Automatic code review using deep multi-instance learning
Heng-Yi Li, Shu-Ting Shi, Ferdian Thung, Xuan Huo, Bowen Xu, Ming Li, and David Lo · 2019
Earlier work this paper cites.
Deep learning-based vulnerable function detection: A benchmark
Guanjun Lin, Wei Xiao, Jun Zhang, and Yang Xiang · 2019
Earlier work this paper cites.
Software vulnerability discovery via learning multi-domain knowledge bases
Guanjun Lin, Jun Zhang, Wei Luo, Lei Pan, Olivier Y. de Vel, Paul Montague, and Yang Xiang · 2019
Earlier work this paper cites.
Learning to spot and refactor inconsistent method names
Kui Liu, Dongsun Kim, Tegawendé F. Bissyandé, Tae-young Kim, Kisub Kim, Anil Koyuncu, Suntae Kim, and Yves Le Traon · 2019
Earlier work this paper cites.
Tbar: revisiting template-based automated program repair
Kui Liu, Anil Koyuncu, Dongsun Kim, and Tegawendé F. Bissyandé · 2019
Earlier work this paper cites.
Generating commit messages from diffs using pointer-generator network
Qin Liu, Zihe Liu, Hongming Zhu, Hongfei Fan, Bowen Du, and Yu Qian · 2019
Earlier work this paper cites.
Aroma: code recommendation via structural code search
Sifei Luan, Di Yang, Celeste Barnaby, Koushik Sen, and Satish Chandra · 2019
Earlier work this paper cites.
BEARS: an extensible java bug benchmark for automatic program repair studies
Fernanda Madeiral, Simon Urli, Marcelo de Almeida Maia, and Martin Monperrus · 2019
Earlier work this paper cites.
Nl2type: inferring javascript function types from natural language information
Rabee Sohail Malik, Jibesh Patra, and Michael Pradel · 2019
Earlier work this paper cites.
Neo: A learned query optimizer
Ryan Marcus, Parimarjan Negi, Hongzi Mao, Chi Zhang, Mohammad Alizadeh, Tim Kraska, Olga Papaemmanouil, and Nesime Tatbul · 2019
Earlier work this paper cites.
Loganomaly: Unsupervised detection of sequential and quantitative anomalies in unstructured logs
Weibin Meng, Ying Liu, Yichen Zhu, Shenglin Zhang, Dan Pei, Yuqing Liu, Yihao Chen, Ruizhi Zhang, Shimin Tao, Pei Sun, and Rong Zhou · 2019
Earlier work this paper cites.
CLCDSA: cross language code clone detection using syntactical features and API documentation
Kawser Wazed Nafi, Tonny Shekha Kar, Banani Roy, Chanchal K. Roy, and Kevin A. Schneider · 2019
Earlier work this paper cites.
Graph-based mining of in-the-wild, fine-grained, semantic code change patterns
Hoan Anh Nguyen, Tien N. Nguyen, Danny Dig, Son Nguyen, Hieu Tran, and Michael Hilton · 2019
Earlier work this paper cites.
Tensorfuzz: Debugging neural networks with coverage-guided fuzzing
Augustus Odena, Catherine Olsson, David G. Andersen, and Ian J. Goodfellow · 2019
Earlier work this paper cites.
Cross-language clone detection by learning over abstract syntax trees
Daniel Perez and Shigeru Chiba · 2019
Earlier work this paper cites.
A manually-curated dataset of fixes to vulnerabilities of open-source software
Serena Elisa Ponta, Henrik Plate, Antonino Sabetta, Michele Bezzi, and Cédric Dangremont · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
NEUZZ: efficient fuzzing with neural program smoothing
Dongdong She, Kexin Pei, Dave Epstein, Junfeng Yang, Baishakhi Ray, and Suman Jana · 2019
Earlier work this paper cites.
Automatic code review by learning the revision of source code
Shu-Ting Shi, Ming Li, David Lo, Ferdian Thung, and Xuan Huo · 2019
Earlier work this paper cites.
Pythia: Ai-assisted code completion system
Alexey Svyatkovskiy, Ying Zhao, Shengyu Fu, and Neel Sundaresan · 2019
Earlier work this paper cites.
An approach to identify use case scenarios from textual requirements specification
Saurabh Tiwari, Deepti Ameta, and Asim Banerjee · 2019
Earlier work this paper cites.
Bugswarm: mining and continuously growing a dataset of reproducible failures and fixes
David A. Tomassi, Naji Dmeiri, Yichen Wang, Antara Bhowmick, Yen-Chuan Liu, Premkumar T. Devanbu, Bogdan Vasilescu, and Cindy Rubio-González · 2019
Earlier work this paper cites.
Recovering variable names for minified code with usage contexts
Hieu Tran, Ngoc M. Tran, Son Nguyen, Hoan Nguyen, and Tien N. Nguyen · 2019
Earlier work this paper cites.
On learning meaningful code changes via neural machine translation
Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk · 2019
Earlier work this paper cites.
Learning how to mutate source code from bug-fixes
Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk · 2019
Earlier work this paper cites.
Neural program repair by jointly learning to localize and repair
Marko Vasic, Aditya Kanade, Petros Maniatis, David Bieber, and Rishabh Singh · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Code generation as a dual task of code summarization
Bolin Wei, Ge Li, Xin Xia, Zhiyi Fu, and Zhi Jin · 2019
Earlier work this paper cites.
Commit message generation for source code changes
Shengbin Xu, Yuan Yao, Feng Xu, Tianxiao Gu, Hanghang Tong, and Jian Lu · 2019
Earlier work this paper cites.
Method name suggestion with hierarchical attention networks
Sihan Xu, Sen Zhang, Weijing Wang, Xinya Cao, Chenkai Guo, and Jing Xu · 2019
Earlier work this paper cites.
Neural detection of semantic code clones via tree-based convolution
Hao Yu, Wing Lam, Long Chen, Ge Li, Tao Xie, and Qianxiang Wang · 2019
Earlier work this paper cites.
Sparc: Cross-domain semantic parsing in context
Tao Yu, Rui Zhang, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li, Heyang Er, Irene Li, Bo Pang, Tao Chen, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Vincent Zhang, Caiming Xiong, Richard Socher, and Dragomir R. Radev · 2019
Earlier work this paper cites.
A novel neural source code representation based on abstract syntax tree
Jian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun, Kaixuan Wang, and Xudong Liu · 2019
Cited alongside, same era.
Editing-based SQL query generation for cross-domain context-dependent questions
Rui Zhang, Tao Yu, Heyang Er, Sungrok Shim, Eric Xue, Xi Victoria Lin, Tianze Shi, Caiming Xiong, Richard Socher, and Dragomir R. Radev · 2019
Cited alongside, same era.
Robust log-based anomaly detection on unstable log data
Xu Zhang, Yong Xu, Qingwei Lin, Bo Qiao, Hongyu Zhang, Yingnong Dang, Chunyu Xie, Xinsheng Yang, Qian Cheng, Ze Li, Junjie Chen, Xiaoting He, Randolph Yao, Jian-Guang Lou, Murali Chintalapati, Furao Shen, and Dongmei Zhang · 2019
Cited alongside, same era.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Yaqin Zhou, Shangqing Liu, Jing Kai Siow, Xiaoning Du, and Yang Liu · 2019
Cited alongside, same era.
Tools and benchmarks for automated log parsing
Jieming Zhu, Shilin He, Jinyang Liu, Pinjia He, Qi Xie, Zibin Zheng, and Michael R. Lyu · 2019
Cited alongside, same era.
BLOOM: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, and et al · 2022
Later among the works it cites.
Self-supervised vision transformers for malware detection
Sachith Seneviratne, Ridwan Shariffdeen, Sanka Rasnayaka, and Nuran Kasthuriarachchi · 2022
Later among the works it cites.
LAMNER: code comment generation using character language model and named entity recognition
Rishab Sharma, Fuxiang Chen, and Fatemeh H. Fard · 2022
Later among the works it cites.
RACE: retrieval-augmented commit message generation
Ensheng Shi, Yanlin Wang, Wei Tao, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
$\mu$ μ \mu vuldeepecker: A deep learning-based system for multiclass vulnerability detection
Deqing Zou, Sujuan Wang, Shouhuai Xu, Zhen Li, and Hai Jin · 2019
Cited alongside, same era.
Generating uml class diagram using nlp techniques and heuristic rules
Esra A. Abdelnabi, Abdelsalam M. Maatuk, Tawfig M. Abdelaziz, and Salwa M. Elakeili · 2020
Cited alongside, same era.
A transformer-based approach for source code summarization
Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang · 2020
Cited alongside, same era.
Typilus: neural type hints
Miltiadis Allamanis, Earl T. Barr, Soline Ducousso, and Zheng Gao · 2020
Cited alongside, same era.
Structural language models of code
Uri Alon, Roy Sadaka, Omer Levy, and Eran Yahav · 2020
Cited alongside, same era.
Learning to recommend third-party library migration opportunities at the API level
Hussein Alrubaye, Mohamed Wiem Mkaouer, Igor Khokhlov, Leon Reznik, Ali Ouni, and Jason Mcgoff · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Leveraging code-test co-evolution patterns for automated test case recommendation
Samiha Shimmi and Mona Rahimi · 2022
Later among the works it cites.
Mining idioms in the wild
Aishwarya Sivaraman, Rui Abreu, Andrew Scott, Tobi Akomolede, and Satish Chandra · 2022
Later among the works it cites.
Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model
Saleh Soltan, Shankar Ananthakrishnan, Jack FitzGerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith Peris, Stephen Rawls, Andy Rosenbaum, Anna Rumshisky, Chandana Satya Prakash, Mukund Sridhar, Fabian Triefenbach, Apurv Verma, Gökhan Tür, and Prem Natarajan · 2022
Later among the works it cites.
Code search based on context-aware code translation
Weisong Sun, Chunrong Fang, Yuchen Chen, Guanhong Tao, Tingxu Han, and Quanjun Zhang · 2022
Later among the works it cites.
Ast-trans: Code summarization with efficient tree-structured attention
Ze Tang, Xiaoyu Shen, Chuanyi Li, Jidong Ge, Liguo Huang, Zheling Zhu, and Bin Luo · 2022
Later among the works it cites.
Logstamp: Automatic online log parsing based on sequence labelling
Shimin Tao, Weibin Meng, Yimeng Chen, Yichen Zhu, Ying Liu, Chunning Du, Tao Han, Yongpeng Zhao, Xiangguang Wang, and Hao Yang · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen S. Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed H. Chi, and Quoc Le · 2022
Later among the works it cites.
Codexdb: Synthesizing code for query processing from natural language instructions using GPT-3 codex
Immanuel Trummer · 2022
Later among the works it cites.
Malicious source code detection using transformer
Chen Tsfaty and Michael Fire · 2022
Later among the works it cites.
Generating accurate assert statements for unit test cases using pretrained transformers
Michele Tufano, Dawn Drain, Alexey Svyatkovskiy, and Neel Sundaresan · 2022
Later among the works it cites.
Using pre-trained models to boost code review automation
Rosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella, Denys Poshyvanyk, and Gabriele Bavota · 2022
Later among the works it cites.
Natural Language Processing with Transformers: Building Language Applications with Hugging Face
Lewis Tunstall, Leandro von Werra, and Thomas Wolf · 2022
Later among the works it cites.
Bridging pre-trained models and downstream tasks for source code understanding
Deze Wang, Zhouyang Jia, Shanshan Li, Yue Yu, Yun Xiong, Wei Dong, and Xiangke Liao · 2022
Later among the works it cites.
Webformer: The web-page transformer for structure information extraction
Qifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng, Xiaojun Quan, and Dongfang Liu · 2022
Later among the works it cites.
What language model architecture and pretraining objective works best for zero-shot generalization?
Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel · 2022
Later among the works it cites.
Compilable neural code generation with compiler feedback
Xin Wang, Yasheng Wang, Yao Wan, Fei Mi, Yitong Li, Pingyi Zhou, Jin Liu, Hao Wu, Xin Jiang, and Qun Liu · 2022
Later among the works it cites.
CODE-MVP: learning to represent source code from multiple views with contrastive pre-training
Xin Wang, Yasheng Wang, Yao Wan, Jiawei Wang, Pingyi Zhou, Li Li, Hao Wu, and Jin Liu · 2022
Later among the works it cites.
SPINE: a scalable log parser with feedback guidance
Xuheng Wang, Xu Zhang, Liqun Li, Shilin He, Hongyu Zhang, Yudong Liu, Lingling Zheng, Yu Kang, Qingwei Lin, Yingnong Dang, Saravanakumar Rajmohan, and Dongmei Zhang · 2022
Later among the works it cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ NLP tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen · 2022
Later among the works it cites.
Free lunch for testing: Fuzzing deep-learning libraries from open source
Anjiang Wei, Yinlin Deng, Chenyuan Yang, and Lingming Zhang · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Later among the works it cites.
Babeltower: Learning to auto-parallelized program translation
Yuanbo Wen, Qi Guo, Qiang Fu, Xiaqing Li, Jianxing Xu, Yanlin Tang, Yongwei Zhao, Xing Hu, Zidong Du, Ling Li, Chao Wang, Xuehai Zhou, and Yunji Chen · 2022
Later among the works it cites.
Recommending metamodel concepts during modeling activities with pre-trained language models
Martin Weyssow, Houari A. Sahraoui, and Eugene Syriani · 2022
Later among the works it cites.
Evaluating and improving neural program-smoothing-based fuzzing
Mingyuan Wu, Ling Jiang, Jiahong Xiang, Yuqun Zhang, Guowei Yang, Huixin Ma, Sen Nie, Shi Wu, Heming Cui, and Lingming Zhang · 2022
Later among the works it cites.
Less training, more repairing please: revisiting automated program repair via zero-shot learning
Chunqiu Steven Xia and Lingming Zhang · 2022
Later among the works it cites.
Docter: documentation-guided fuzzing for testing deep learning API functions
Danning Xie, Yitong Li, Mijung Kim, Hung Viet Pham, Lin Tan, Xiangyu Zhang, and Michael W. Godfrey · 2022
Later among the works it cites.
Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models
Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, and Tao Yu · 2022
Later among the works it cites.
A systematic evaluation of large language models of code
Frank F. Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn · 2022
Later among the works it cites.
A survey on pretrained language models for neural code intelligence
Yichen Xu and Yanqiao Zhu · 2022
Later among the works it cites.
Cross-language source code clone detection using deep learning with infercode
Mohammad A. Yahya and Dae-Kyoo Kim · 2022
Later among the works it cites.
A naming pattern based approach for method name recommendation
Yanping Yang, Ling Xu, Meng Yan, Zhou Xu, and Zhongyang Deng · 2022
Later among the works it cites.
CERT: continual pre-training on sketches for library-oriented code generation
Daoguang Zan, Bei Chen, Dejian Yang, Zeqi Lin, Minsu Kim, Bei Guan, Yongji Wang, Weizhu Chen, and Jian-Guang Lou · 2022
Later among the works it cites.
A survey of automatic source code summarization
Chunyan Zhang, Junchao Wang, Qinglei Zhou, Ting Xu, Ke Tang, Hairen Gui, and Fudong Liu · 2022
Later among the works it cites.
System log parsing: A survey
Tianzhu Zhang, Han Qiu, Gabriele Castellano, Myriana Rifai, Chung Shue Chen, and Fabio Pianese · 2022
Later among the works it cites.
Neural program repair : Systems, challenges and solutions
Wenkang Zhong, Chuanyi Li, Jidong Ge, and Bin Luo · 2022
Later among the works it cites.
Multilingual code snippets training for program translation
Ming Zhu, Karthik Suresh, and Chandan K. Reddy · 2022
Later among the works it cites.
Sallam Abualhaija, Marcello Ceci, and Lionel C. Briand · 2023
Closest in time.
Guiding language models of code with global context using monitors
Lakshya A Agrawal, Aditya Kanade, Navin Goyal, Shuvendu K. Lahiri, and Sriram K. Rajamani · 2023
Closest in time.
AVATAR: A parallel corpus for java-python program translation
Wasi Uddin Ahmad, Md Golam Rahman Tushar, Saikat Chakraborty, and Kai-Wei Chang · 2023
Closest in time.
GQA: training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Closest in time.
A3test: Assertion-augmented automated test case generation
Saranya Alagarsamy, Chakkrit Tantithamthavorn, and Aldeida Aleti · 2023
Closest in time.
Santacoder: don’t reach for the stars!
Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Muñoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, Logesh Kumar Umapathi, Carolyn Jane Anderson, Yangtian Zi, Joel Lamy-Poirier, Hailey Schoelkopf, Sergey Troshin, Dmitry Abulkhanov, Manuel Romero, Michael Lappert, Francesco De Toni, Bernardo García del Río, Qian Liu, Shamik Bose, Urvashi Bhattacharyya, Terry Yue Zhuo, Ian Yu, Paulo Villegas, Marco Zocca, Sourab Mangrulkar, David Lansky, Huu Nguyen, Danish Contractor, Luis Villa, Jia Li, Dzmitry Bahdanau, Yacine Jernite, Sean Hughes, Daniel Fried, Arjun Guha, Harm de Vries, and Leandro von Werra · 2023
Closest in time.
Slade: A portable small language model decompiler for optimized assembler
Jordi Armengol-Estapé, Jackson Woodruff, Chris Cummins, and Michael F. P. O’Boyle · 2023
Closest in time.
Advancing requirements engineering through generative AI: assessing the role of llms
Chetan Arora, John Grundy, and Mohamed Abdelrazek · 2023
Closest in time.
Is github’s copilot as bad as humans at introducing vulnerabilities in code?
Owura Asare, Meiyappan Nagappan, and N. Asokan · 2023
Closest in time.
Multi-lingual evaluation of code generation models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li, Yuchen Tian, Ming Tan, Wasi Uddin Ahmad, Shiqi Wang, Qing Sun, Mingyue Shang, Sujan Kumar Gonugondla, Hantian Ding, Varun Kumar, Nathan Fulton, Arash Farahani, Siddhartha Jain, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, and Ramesh Nallapati · 2023
Closest in time.
Studenteval: A benchmark of student-written prompts for large language models of code
Hannah McLean Babe, Sydney Nguyen, Yangtian Zi, Arjun Guha, Molly Q. Feldman, and Carolyn Jane Anderson · 2023
Closest in time.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu · 2023
Closest in time.
Codeplan: Repository-level coding using llms and planning
Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D. C, Arun Iyer, Suresh Parthasarathy, Sriram K. Rajamani, Balasubramanyan Ashok, and Shashank Shet · 2023
Closest in time.
Property-based mutation testing
Ezio Bartocci, Leonardo Mariani, Dejan Nickovic, and Drishti Yadav · 2023
Closest in time.
Benchmarking software vulnerability detection techniques: A survey
Yingzhou Bi, Jiangtao Huang, Penghui Liu, and Lianmei Wang · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco Túlio Ribeiro, and Yi Zhang · 2023
Closest in time.
On the assessment of generative AI in modeling tasks: an experience report with chatgpt and UML
Javier Cámara, Javier Troya, Lola Burgueño, and Antonio Vallecillo · 2023
Closest in time.
A study on prompt design, advantages and limitations of chatgpt for deep learning program repair
Jialun Cao, Meiziniu Li, Ming Wen, and Shing-Chi Cheung · 2023
Closest in time.
Towards using few-shot prompt learning for automating model completion
Meriem Ben Chaaben, Lola Burgueño, and Houari A. Sahraoui · 2023
Closest in time.
Ernie-code: Beyond english-centric cross-lingual pretraining for programming languages
Yekun Chai, Shuohuan Wang, Chao Pang, Yu Sun, Hao Tian, and Hua Wu · 2023
Closest in time.
Transformer-based vulnerability detection in code at edittime: Zero-shot, few-shot, or fine-tuning?
Aaron Chan, Anant Kharkar, Roshanak Zilouchian Moghaddam, Yevhen Mohylevskyy, Alec Helyar, Eslam Kamal, Mohamed Elkamhawy, and Neel Sundaresan · 2023
Closest in time.
How to prompt llms for text-to-sql: A study in zero-shot, single-domain, and cross-domain settings
Shuaichen Chang and Eric Fosler-Lussier · 2023
Closest in time.
Codet: Code generation with generated tests
Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen · 2023
Closest in time.
On the use of GPT-4 for creating goal models: An exploratory study
Boqi Chen, Kua Chen, Shabnam Hassani, Yujing Yang, Daniel Amyot, Lysanne Lessard, Gunter Mussbacher, Mehrdad Sabetzadeh, and Dániel Varró · 2023
Closest in time.
Automated domain modeling with large language models: A comparative study
Kua Chen, Yujing Yang, Boqi Chen, José Antonio Hernández López, Gunter Mussbacher, and Dániel Varró · 2023
Closest in time.
Diversevul: A new vulnerable source code dataset for deep learning based vulnerability detection
Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David A. Wagner · 2023
Closest in time.
Play: Parametrically conditioned layout generation using latent diffusion
Chin-Yi Cheng, Forrest Huang, Gang Li, and Yang Li · 2023
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2023
Closest in time.
Large language models for compiler optimization
Chris Cummins, Volker Seeker, Dejan Grubisic, Mostafa Elhoushi, Youwei Liang, Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Kim M. Hazelwood, Gabriel Synnaeve, and Hugh Leather · 2023
Closest in time.
Effective test generation using pre-trained large language models and mutation testing
Arghavan Moradi Dakhel, Amin Nikanjam, Vahid Majdinasab, Foutse Khomh, and Michel C. Desmarais · 2023
Closest in time.
Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models
Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang · 2023
Closest in time.
Codefuse-13b: A pretrained multi-lingual code large language model
Peng Di, Jianguo Li, Hang Yu, Wei Jiang, Wenting Cai, Yang Cao, Chaoyu Chen, Dajun Chen, Hongwei Chen, Liang Chen, Gang Fan, Jie Gong, Zi Gong, Wen Hu, Tingting Guo, Zhichao Lei, Ting Li, Zheng Li, Ming Liang, Cong Liao, Bingchang Liu, Jiachen Liu, Zhiwei Liu, Shaojun Lu, Min Shen, Guangpei Wang, Huan Wang, Zhi Wang, Zhaogui Xu, Jiawei Yang, Qing Ye, Gehao Zhang, Yu Zhang, Zelin Zhao, Xunjin Zheng, Hailian Zhou, Lifu Zhu, and Xianying Zhu · 2023
Closest in time.
Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion
Yangruibo Ding, Zijian Wang, Wasi Uddin Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang · 2023
Closest in time.
Self-collaboration code generation via chatgpt
Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li · 2023
Closest in time.
Towards understanding the capability of large language models on code clone detection: A survey
Shihan Dou, Junjie Shan, Haoxiang Jia, Wenhao Deng, Zhiheng Xi, Wei He, Yueming Wu, Tao Gui, Yang Liu, and Xuanjing Huang · 2023
Closest in time.
A comparative analysis of large language models for code documentation generation
Shubhang Shekhar Dvivedi, Vyshnav Vijay, Sai Leela Rahul Pujari, Shoumik Lodh, and Dhruv Kumar · 2023
Closest in time.
From commit message generation to history-aware commit message completion
Aleksandra Eliseeva, Yaroslav Sokolov, Egor Bogomolov, Yaroslav Golubev, Danny Dig, and Timofey Bryksin · 2023
Closest in time.
Automated repair of programs from large language models
Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan · 2023
Closest in time.
Layoutgpt: Compositional visual planning and generation with large language models
Weixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani, Arjun R. Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang · 2023
Closest in time.
Incoder: A generative model for code infilling and synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis · 2023
Closest in time.
Codeapex: A bilingual programming evaluation benchmark for large language models
Lingyue Fu, Huacan Chai, Shuang Luo, Kounianhua Du, Weiming Zhang, Longteng Fan, Jiayi Lei, Renting Rui, Jianghao Lin, Yuchen Fang, Yifan Liu, Jingkuan Wang, Siyuan Qi, Kangning Zhang, Weinan Zhang, and Yong Yu · 2023
Closest in time.
PAL: program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Closest in time.
Code search: A survey of techniques for finding code
Luca Di Grazia and Michael Pradel · 2023
Closest in time.
An initial investigation of chatgpt unit test generation capability
Vitor Guilherme and Auri Vincenzi · 2023
Closest in time.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li · 2023
Closest in time.
Retrieval-augmented gpt-3.5-based text-to-sql framework with sample-aware prompting and dynamic revision chain
Chunxi Guo, Zhiliang Tian, Jintao Tang, Shasha Li, Zhihua Wen, Kaixuan Wang, and Ting Wang · 2023
Closest in time.
Longcoder: A long-range pre-trained language model for code completion
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian J. McAuley · 2023
Closest in time.
Understanding HTML with large language models
Izzeddin Gur, Ofir Nachum, Yingjie Miao, Mustafa Safdari, Austin Huang, Aakanksha Chowdhery, Sharan Narang, Noah Fiedel, and Aleksandra Faust · 2023
Closest in time.
A survey on automated software vulnerability detection using machine learning and deep learning
Nima Shiri Harzevili, Alvine Boaye Belle, Junjie Wang, Song Wang, Zhen Ming Jiang, and Nachiappan Nagappan · 2023
Closest in time.
COME: commit message generation with modification embedding
Yichen He, Liran Wang, Kaiyi Wang, Yupeng Zhang, Hang Zhang, and Zhoujun Li · 2023
Closest in time.
Solving math word problems by combining language models with symbolic solvers
Joy He-Yueya, Gabriel Poesia, Rose E. Wang, and Noah D. Goodman · 2023
Closest in time.
Metagpt: Meta programming for multi-agent collaborative framework
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, and Chenglin Wu · 2023
Closest in time.
Unnatural instructions: Tuning language models with (almost) no human labor
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick · 2023
Closest in time.
Large language models for software engineering: A systematic literature review
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John C. Grundy, and Haoyu Wang · 2023
Closest in time.
A transformer-based approach for abstractive summarization of requirements from obligations in software engineering contracts
Chirag Jain, Preethu Rose Anish, Amrita Singh, and Smita Ghaisas · 2023
Closest in time.
Attention, compilation, and solver-based symbolic analysis are all you need
Prithwish Jana, Piyush Jha, Haoyang Ju, Gautham Kishore, Aryan Mahajan, and Vijay Ganesh · 2023
Closest in time.
Large language models and simple, stupid bugs
Kevin Jesse, Toufique Ahmed, Premkumar T. Devanbu, and Emily Morgan · 2023
Closest in time.
Impact of code language models on automated program repair
Nan Jiang, Kevin Liu, Thibaud Lutellier, and Lin Tan · 2023
Closest in time.
On the evaluation of neural code translation: Taxonomy and benchmark
Mingsheng Jiao, Tingrui Yu, Xuan Li, Guanjie Qiu, Xiaodong Gu, and Beijun Shen · 2023
Closest in time.
Swe-bench: Can language models resolve real-world github issues?
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Closest in time.
Repair is nearly generation: Multilingual program repair with llms
Harshit Joshi, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le, Gust Verbruggen, and Ivan Radicek · 2023
Closest in time.
Menucraft: Interactive menu system design with large language models
Amir Hossein Kargaran, Nafiseh Nikeghbal, Abbas Heydarnoori, and Hinrich Schütze · 2023
Closest in time.
A survey on deep learning approaches for text-to-sql
George Katsogiannis-Meimarakis and Georgia Koutrika · 2023
Closest in time.
Mohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Long Do, Weishi Wang, Md. Rizwan Parvez, and Shafiq R. Joty · 2023
Closest in time.
The stack: 3 TB of permissively licensed source code
Denis Kocetkov, Raymond Li, Loubna Ben allal, Jia LI, Chenghao Mou, Yacine Jernite, Margaret Mitchell, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro Von Werra, and Harm de Vries · 2023
Closest in time.
A word is worth a thousand pictures: Prompts as AI design material
Chinmay Kulkarni, Stefania Druga, Minsuk Chang, Alex Fiannaca, Carrie J. Cai, and Michael Terry · 2023
Closest in time.
DS-1000: A natural and reliable benchmark for data science code generation
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-Tau Yih, Daniel Fried, Sida I. Wang, and Tao Yu · 2023
Closest in time.
Log parsing with prompt-based few-shot learning
Van-Hoang Le and Hongyu Zhang · 2023
Closest in time.
Log parsing: How far can chatgpt go?
Van-Hoang Le and Hongyu Zhang · 2023
Closest in time.
Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models
Caroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, and Siddhartha Sen · 2023
Closest in time.
RESDSQL: decoupling schema linking and skeleton parsing for text-to-sql
Haoyang Li, Jing Zhang, Cuiping Li, and Hong Chen · 2023
Closest in time.
Graphix-t5: Mixing pre-trained transformers with graph-aware layers for text-to-sql parsing
Jinyang Li, Binyuan Hui, Reynold Cheng, Bowen Qin, Chenhao Ma, Nan Huo, Fei Huang, Wenyu Du, Luo Si, and Yongbin Li · 2023
Closest in time.
Dianshu Liao, Shidong Pan, Xiaoyu Sun, Xiaoxue Ren, Qing Huang, Zhenchang Xing, Huan Jin, and Qinying Li · 2023
Closest in time.
Text generation with diffusion language models: A pre-training approach with continuous paragraph denoise
Zhenghao Lin, Yeyun Gong, Yelong Shen, Tong Wu, Zhihao Fan, Chen Lin, Nan Duan, and Weizhu Chen · 2023
Closest in time.
Syntax and domain aware model for unsupervised program translation
Fang Liu, Jia Li, and Li Zhang · 2023
Closest in time.
Nnsmith: Generating diverse and valid test cases for deep learning compilers
Jiawei Liu, Jinkun Lin, Fabian Ruffy, Cheng Tan, Jinyang Li, Aurojit Panda, and Lingming Zhang · 2023
Closest in time.
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Closest in time.
On ml-based program translation: Perils and promises
Aniketh Malyala, Katelyn Zhou, Baishakhi Ray, and Saikat Chakraborty · 2023
Closest in time.
Automated devops pipeline generation for code repositories using large language models
Deep Mehta, Kartik Rawool, Subodh Gujar, and Bowen Xu · 2023
Closest in time.
Developer-intent driven code comment generation
Fangwen Mu, Xiao Chen, Lin Shi, Song Wang, and Qing Wang · 2023
Closest in time.
An assessment of chatgpt on log data
Priyanka Mudgal and Rita H. Wouhaybi · 2023
Closest in time.
Linyong Nan, Yilun Zhao, Weijin Zou, Narutatsu Ri, Jaesung Tae, Ellen Zhang, Arman Cohan, and Dragomir Radev · 2023
Closest in time.
Learning deep semantics for test completion
Pengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney, and Milos Gligoric · 2023
Closest in time.
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2023
Closest in time.
An empirical comparison of pre-trained models of source code
Changan Niu, Chuanyi Li, Vincent Ng, Dongxiao Chen, Jidong Ge, and Bin Luo · 2023
Closest in time.
On comparing mutation testing tools through learning-based mutant selection
Milos Ojdanic, Ahmed Khanfir, Aayush Garg, Renzo Degiovanni, Mike Papadakis, and Yves Le Traon · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Understanding the effectiveness of large language models in code translation
Rangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar, Lambert Pouguem Wassi, Michele Merler, Boris Sobolev, Raju Pavuluri, Saurabh Sinha, and Reyhaneh Jabbarvand · 2023
Closest in time.
Automated program repair based on code review: How do pre-trained transformer models perform?
Rishov Paul, Md. Mohib Hossain, Masum Hasan, and Anindya Iqbal · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Closest in time.
Do users write more insecure code with AI assistants?
Neil Perry, Megha Srivastava, Deepak Kumar, and Dan Boneh · 2023
Closest in time.
DIN-SQL: decomposed in-context learning of text-to-sql with self-correction
Mohammadreza Pourreza and Davood Rafiei · 2023
Closest in time.
Runbugrun - an executable dataset for automated program repair
Julian Aron Prenner and Romain Robbes · 2023
Closest in time.
Loggpt: Exploring chatgpt for log-based anomaly detection
Jiaxing Qi, Shaohan Huang, Zhongzhi Luan, Carol J. Fung, Hailong Yang, and Depei Qian · 2023
Closest in time.
Communicative agents for software development
Chen Qian, Xin Cong, Cheng Yang, Weize Chen, Yusheng Su, Juyuan Xu, Zhiyuan Liu, and Maosong Sun · 2023
Closest in time.
Malbertv2: Code aware bert-based model for malware identification
Abir Rahali and Moulay A. Akhloufi · 2023
Closest in time.
Towards causal deep learning for vulnerability detection
Md Mahbubur Rahman, Ira Ceka, Chengzhi Mao, Saikat Chakraborty, Baishakhi Ray, and Wei Le · 2023
Closest in time.
Limits of machine learning for automatic vulnerability detection
Niklas Risse and Marcel Böhme · 2023
Closest in time.
Prompts matter: Insights and strategies for prompt engineering in automated software traceability
Alberto D. Rodriguez, Katherine R. Dearstyne, and Jane Cleland-Huang · 2023
Closest in time.
Requirements engineering using generative AI: prompts and prompting patterns
Krishna Ronanki, Beatriz Cabrero Daniel, Jennifer Horkoff, and Christian Berger · 2023
Closest in time.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve · 2023
Closest in time.
TPTU: task planning and tool usage of large language model-based AI agents
Jingqing Ruan, Yihong Chen, Bin Zhang, Zhiwei Xu, Tianpeng Bao, Guoqing Du, Shiwei Shi, Hangyu Mao, Xingyu Zeng, and Rui Zhao · 2023
Closest in time.
On contrastive learning of semantic similarity forcode to code search
Anthony Saieva, Saikat Chakraborty, and Gail E. Kaiser · 2023
Closest in time.
Lost at C: A user study on the security implications of large language model code assistants
Gustavo Sandoval, Hammond Pearce, Teo Nys, Ramesh Karri, Siddharth Garg, and Brendan Dolan-Gavitt · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Closest in time.
An empirical evaluation of using large language models for automated unit test generation, 2023
Max Schäfer, Sarah Nadi, Aryaz Eghbali, and Frank Tip · 2023
Closest in time.
Apicontext2com: Code comment generation by incorporating pre-defined API documentation
Ramin Shahbazi and Fatemeh Hendijani Fard · 2023
Closest in time.
Pitfalls in language models for code intelligence: A taxonomy and survey
Xinyu She, Yue Liu, Yanjie Zhao, Yiling He, Li Li, Chakkrit Tantithamthavorn, Zhan Qin, and Haoyu Wang · 2023
Closest in time.
Pangu-coder2: Boosting large language models for code with ranking feedback
Bo Shen, Jiaxin Zhang, Taihong Chen, Daoguang Zan, Bing Geng, An Fu, Muhan Zeng, Ailun Yu, Jichuan Ji, Jingyang Zhao, Yuenan Guo, and Qianxiang Wang · 2023
Closest in time.
Coss: Leveraging statement semantics for code summarization
Chaochen Shi, Borui Cai, Yao Zhao, Longxiang Gao, Keshav Sood, and Yong Xiang · 2023
Closest in time.
Cocosoda: Effective contrastive learning for code search
Ensheng Shi, Yanlin Wang, Wenchao Gu, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun · 2023
Closest in time.
Domain adaptation for deep unit test case generation
Jiho Shin, Sepehr Hashtroudi, Hadi Hemmati, and Song Wang · 2023
Closest in time.
Execution-based code generation using deep reinforcement learning
Parshin Shojaee, Aneesh Jain, Sindhu Tipirneni, and Chandan K. Reddy · 2023
Closest in time.
Repository-level prompt generation for large language models of code
Disha Shrivastava, Hugo Larochelle, and Daniel Tarlow · 2023
Closest in time.
Codefusion: A pre-trained diffusion model for code generation, 2023
Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le, Carina Negreanu, and Gust Verbruggen · 2023
Closest in time.
Cct-code: Cross-consistency training for multilingual clone detection and code search
Nikita Sorokin, Dmitry Abulkhanov, Sergey I. Nikolenko, and Valentin Malykh · 2023
Closest in time.
Learning ui-to-code reverse generator using visual critic without rendering
Davit Soselia, Khalid Saifullah, and Tianyi Zhou · 2023
Closest in time.
An empirical study of deep learning models for vulnerability detection
Benjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, and Wei Le · 2023
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha H. M. Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2023
Closest in time.
Hotgpt: How to make software documentation more useful with a large language model?
Yiming Su, Chengcheng Wan, Utsav Sethi, Shan Lu, Madan Musuvathi, and Suman Nath · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick · 2023
Closest in time.
Code translation with compiler representations
Marc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut, François Charton, and Gabriel Synnaeve · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
UL2: unifying language learning paradigms
Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler · 2023
Closest in time.
Benchmarking large language models for automated verilog RTL code generation
Shailja Thakur, Baleegh Ahmad, Zhenxing Fan, Hammond Pearce, Benjamin Tan, Ramesh Karri, Brendan Dolan-Gavitt, and Siddharth Garg · 2023
Closest in time.
Kiran Thorat, Jiahui Zhao, Yaotian Liu, Hongwu Peng, Xi Xie, Bin Lei, Jeff Zhang, and Caiwen Ding · 2023
Closest in time.
Rtlfixer: Automatically fixing RTL syntax errors with large language models
Yun-Da Tsai, Mingjie Liu, and Haoxing Ren · 2023
Closest in time.
Can large language models identify and reason about security vulnerabilities? not yet
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse K. Coskun, and Gianluca Stringhini · 2023
Closest in time.
Can large language models write good property-based tests?
Vasudev Vikram, Caroline Lemieux, and Rohan Padhye · 2023
Closest in time.
Enabling conversational interaction with mobile UI using large language models
Bryan Wang, Gang Li, and Yang Li · 2023
Closest in time.
Delving into commit-issue correlation to enhance commit message generation models
Liran Wang, Xunzhu Tang, Yichen He, Changyu Ren, Shuhua Shi, Chaoran Yan, and Zhoujun Li · 2023
Closest in time.
Designing responsible AI: adaptations of UX practice to meet responsible AI challenges
Qiaosi Wang, Michael Madaio, Shaun K. Kane, Shivani Kapania, Michael Terry, and Lauren Wilcox · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
Mconala: A benchmark for code generation from multiple natural languages
Zhiruo Wang, Grace Cuenca, Shuyan Zhou, Frank F. Xu, and Graham Neubig · 2023
Closest in time.
Typet5: Seq2seq type inference using static analysis
Jiayi Wei, Greg Durrett, and Isil Dillig · 2023
Closest in time.
Jules White, Sam Hays, Quchen Fu, Jesse Spencer-Smith, and Douglas C. Schmidt · 2023
Closest in time.
Refining decompiled C code with large language models
Wai Kin Wong, Huaijin Wang, Zongjie Li, Zhibo Liu, Shuai Wang, Qiyi Tang, Sen Nie, and Shi Wu · 2023
Closest in time.
Conversational automated program repair
Chunqiu Steven Xia and Lingming Zhang · 2023
Closest in time.
Automated program repair in the era of large pre-trained language models
Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang · 2023
Closest in time.
Lmpa: Improving decompilation by synergy of large language model and program analysis
Xiangzhe Xu, Zhuo Zhang, Shiwei Feng, Yapeng Ye, Zian Su, Nan Jiang, Siyuan Cheng, Lin Tan, and Xiangyu Zhang · 2023
Closest in time.
Codetransocean: A comprehensive multilingual benchmark for code translation
Weixiang Yan, Yuchen Tian, Yunzhe Li, Qian Chen, and Wen Wang · 2023
Closest in time.
Fuzzing automatic differentiation in deep-learning libraries
Chenyuan Yang, Yinlin Deng, Jiayi Yao, Yuxing Tu, Hanchi Li, and Lingming Zhang · 2023
Closest in time.
Intercode: Standardizing and benchmarking interactive coding with execution feedback
John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao · 2023
Closest in time.
Do machine learning models produce typescript types that type check?
Ming-Ho Yee and Arjun Guha · 2023
Closest in time.
Burak Yetistiren, Isik Özsoy, Miray Ayerdem, and Eray Tüzün · 2023
Closest in time.
Automatic code review by learning the structure information of code graph
Ying Yin, Yuhai Zhao, Yiming Sun, and Chen Chen · 2023
Closest in time.
Self-supervised log parsing using semantic contribution difference
Siyu Yu, Ningjiang Chen, Yifan Wu, and Wensheng Dou · 2023
Closest in time.
Large language models meet nl2code: A survey
Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang, and Jian-Guang Lou · 2023
Closest in time.
Self-edit: Fault-aware code editor for code generation
Kechi Zhang, Zhuo Li, Jia Li, Ge Li, and Zhi Jin · 2023
Closest in time.
Length extrapolation of transformers: A survey from the perspective of position encoding
Liang Zhao, Xiaocheng Feng, Xiachong Feng, Bing Qin, and Ting Liu · 2023
Closest in time.
Hybrid API migration: A marriage of small API mapping models and large language models
Bingzhe Zhou, Xinying Wang, Shengbin Xu, Yuan Yao, Minxue Pan, Feng Xu, and Xiaoxing Ma · 2023
Closest in time.
Towards retrieval-based neural code summarization: A meta-learning approach
Ziyi Zhou, Huiqun Yu, Guisheng Fan, Zijie Huang, and Kang Yang · 2023
Closest in time.
Phi-3 technical report: A highly capable language model locally on your phone
Marah I Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat S. Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Caio César Teodoro Mendes, Weizhu Chen, Vishrav Chaudhary, Parul Chopra, Allie Del Giorno, Gustavo de Rosa, Matthew Dixon, Ronen Eldan, Dan Iter, Amit Garg, Abhishek Goswami, Suriya Gunasekar, Emman Haider, Junheng Hao, Russell J. Hewett, Jamie Huynh, Mojan Javaheripi, Xin Jin, Piero Kauffmann, Nikos Karampatziakis, Dongwoo Kim, Mahoud Khademi, Lev Kurilenko, James R. Lee, Yin Tat Lee, Yuanzhi Li, Chen Liang, Weishung Liu, Eric Lin, Zeqi Lin, Piyush Madan, Arindam Mitra, Hardik Modi, Anh Nguyen, Brandon Norick, Barun Patra, Daniel Perez-Becker, Thomas Portet, Reid Pryzant, Heyang Qin, Marko Radmilac, Corby Rosset, Sambudha Roy, Olatunji Ruwase, Olli Saarikivi, Amin Saied, Adil Salim, Michael Santacroce, Shital Shah, Ning Shang, Hiteshi Sharma, Xia Song, Masahiro Tanaka, Xin Wang, Rachel Ward, Guanhua Wang, Philipp Witte, Michael Wyatt, Can Xu, Jiahang Xu, Sonali Yadav, Fan Yang, Ziyi Yang, Donghan Yu, Chengruidong Zhang, Cyril Zhang, Jianwen Zhang, Li Lyna Zhang, Yi Zhang, Yue Zhang, Yunan Zhang, and Xiren Zhou · 2024
Closest in time.
Towards standards-compliant assistive technology product specifications via llms
Chetan Arora, John Grundy, Louise Puli, and Natasha Layton · 2024
Closest in time.
Zhangqian Bi, Yao Wan, Zheng Wang, Hongyu Zhang, Batu Guan, Fangxin Lu, Zili Zhang, Yulei Sui, Xuanhua Shi, and Hai Jin · 2024
Closest in time.
Hyungjoo Chae, Yeonghyeon Kim, Seungone Kim, Kai Tzu iunn Ong, Beong woo Kwak, Moohyeon Kim, Seonghwan Kim, Taeyoon Kwon, Jiwan Chung, Youngjae Yu, and Jinyoung Yeo · 2024
Closest in time.
Comments as natural logic pivots: Improve code generation via comment perspective
Yijie Chen, Yijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu, and Jie Zhou · 2024
Closest in time.
Deepseek llm: Scaling open-source language models with longtermism
DeepSeek-AI, Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, Huazuo Gao, Kaige Gao, Wenjun Gao, Ruiqi Ge, Kang Guan, Daya Guo, Jianzhong Guo, Guangbo Hao, Zhewen Hao, Ying He, Wenjie Hu, Panpan Huang, Erhang Li, Guowei Li, Jiashi Li, Yao Li, Y. K. Li, Wenfeng Liang, Fangyun Lin, A. X. Liu, Bo Liu, Wen Liu, Xiaodong Liu, Xin Liu, Yiyuan Liu, Haoyu Lu, Shanghao Lu, Fuli Luo, Shirong Ma, Xiaotao Nie, Tian Pei, Yishi Piao, Junjie Qiu, Hui Qu, Tongzheng Ren, Zehui Ren, Chong Ruan, Zhangli Sha, Zhihong Shao, Junxiao Song, Xuecheng Su, Jingxiang Sun, Yaofeng Sun, Minghui Tang, Bingxuan Wang, Peiyi Wang, Shiyu Wang, Yaohui Wang, Yongji Wang, Tong Wu, Y. Wu, Xin Xie, Zhenda Xie, Ziwei Xie, Yiliang Xiong, Hanwei Xu, R. X. Xu, Yanhong Xu, Dejian Yang, Yuxiang You, Shuiping Yu, Xingkai Yu, B. Zhang, Haowei Zhang, Lecong Zhang, Liyue Zhang, Mingchuan Zhang, Minghua Zhang, Wentao Zhang, Yichao Zhang, Chenggang Zhao, Yao Zhao, Shangyan Zhou, Shunfeng Zhou, Qihao Zhu, and Yuheng Zou · 2024
Closest in time.
CYCLE: learning to self-refine the code generation
Yangruibo Ding, Marcus J. Min, Gail E. Kaiser, and Baishakhi Ray · 2024
Closest in time.
Stepcoder: Improve code generation with reinforcement learning from compiler feedback
Shihan Dou, Yan Liu, Haoxiang Jia, Limao Xiong, Enyu Zhou, Wei Shen, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou, Tao Ji, Rui Zheng, Qi Zhang, Xuanjing Huang, and Tao Gui · 2024
Closest in time.
Canvil: Designerly adaptation for llm-powered user experiences
K. J. Kevin Feng, Q. Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan, Amy X. Zhang, and David W. McDonald · 2024
Closest in time.
Model generation from requirements with llms: an exploratory study
Alessio Ferrari, Sallam Abualhaija, and Chetan Arora · 2024
Closest in time.
Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning
Mingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang, Ge Li, Zhi Jin, Xiaoguang Mao, and Xiangke Liao · 2024
Closest in time.
AST-T5: structure-aware pretraining for code generation and understanding
Linyuan Gong, Mostafa Elhoushi, and Alvin Cheung · 2024
Closest in time.
Priority sampling of large language models for compilers
Dejan Grubisic, Volker Seeker, Gabriel Synnaeve, Hugh Leather, John M. Mellor-Crummey, and Chris Cummins · 2024
Closest in time.
Cruxeval: A benchmark for code reasoning, understanding and execution
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I. Wang · 2024
Closest in time.
Codeip: A grammar-guided multi-bit watermark for large language models of code
Batu Guan, Yao Wan, Zhangqian Bi, Zheng Wang, Hongyu Zhang, Yulei Sui, Pan Zhou, and Lichao Sun · 2024
Closest in time.
Deepseek-coder: When the large language model meets programming - the rise of code intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang · 2024
Closest in time.
Analyzing the performance of large language models on code summarization
Rajarshi Haldar and Julia Hockenmaier · 2024
Closest in time.
PECC: problem extraction and coding challenges
Patrick Haller, Jonas Golde, and Alan Akbik · 2024
Closest in time.
Sivana Hamer, Marcelo d’Amorim, and Laurie Williams · 2024
Closest in time.
CONLINE: complex code generation and refinement with online searching and correctness testing
Xinyi He, Jiaru Zou, Yun Lin, Mengyu Zhou, Shi Han, Zejian Yuan, and Dongmei Zhang · 2024
Closest in time.
Soap: Enhancing efficiency of generated code via self-optimization
Dong Huang, Jianbo Dai, Han Weng, Puzhen Wu, Yuhao Qing, Jie M Zhang, Heming Cui, and Zhijiang Guo · 2024
Closest in time.
Yoichi Ishibashi and Yoshimasa Nishimura · 2024
Closest in time.
Unlocking the conversion of web screenshots into HTML code with the websight dataset
Hugo Laurençon, Léo Tronchon, and Victor Sanh · 2024
Closest in time.
When llm-based code generation meets the software development process
Feng Lin, Dong Jae Kim, and Tse-Husn Chen · 2024
Closest in time.
Exploring and evaluating hallucinations in llm-powered code generation
Fang Liu, Yang Liu, Lin Shi, Houkun Huang, Ruifeng Wang, Zhen Yang, and Li Zhang · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo, Evgenii Zheltonozhskii, Nii Osae Osae Dade, Wenhao Yu, Lucas Krauß, Naman Jain, Yixuan Su, Xuanli He, Manan Dey, Edoardo Abati, Yekun Chai, Niklas Muennighoff, Xiangru Tang, Muhtasham Oblokulov, Christopher Akiki, Marc Marone, Chenghao Mou, Mayank Mishra, Alex Gu, Binyuan Hui, Tri Dao, Armel Zebaze, Olivier Dehaene, Nicolas Patry, Canwen Xu, Julian J. McAuley, Han Hu, Torsten Scholak, Sébastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, and et al · 2024
Closest in time.
Activation steering for robust type prediction in codellms
Francesca Lucchetti and Arjun Guha · 2024
Closest in time.
Llmparser: An exploratory study on using large language models for log parsing
Zeyang Ma, An Ran Chen, Dong Jae Kim, Tse-Hsun Chen, and Shaowei Wang · 2024
Closest in time.
Marcos Macedo, Yuan Tian, Filipe Roseiro Côgo, and Bram Adams · 2024
Closest in time.
Are human rules necessary? generating reusable apis with cot reasoning and in-context learning
Yubo Mai, Zhipeng Gao, Xing Hu, Lingfeng Bao, Yu Liu, and Jianling Sun · 2024
Closest in time.
Multi-role consensus through llms discussions for vulnerability detection
Zhenyu Mao, Jialong Li, Munan Li, and Kenji Tei · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, and et al · 2024
Closest in time.
Octopack: Instruction tuning code large language models
Niklas Muennighoff, Qian Liu, Armel Randy Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro Von Werra, and Shayne Longpre · 2024
Closest in time.
A multi-expert large language model architecture for verilog code generation
Bardia Nadimi and Hao Zheng · 2024
Closest in time.
Can large language models write parallel code?
Daniel Nichols, Joshua Hoke Davis, Zhaojun Xie, Arjun Rajaram, and Abhinav Bhatele · 2024
Closest in time.
On evaluating the efficiency of source code generated by llms
Changan Niu, Ting Zhang, Chuanyi Li, Bin Luo, and Vincent Ng · 2024
Closest in time.
Csa-trans: Code structure aware transformer for ast
Saeyoon Oh and Shin Yoo · 2024
Closest in time.
Xinglu Pan, Chenxiao Liu, Yanzhen Zou, Tao Xie, and Bing Xie · 2024
Closest in time.
The fact selection problem in llm-based program repair
Nikhil Parasaram, Huijie Yan, Boyu Yang, Zineb Flahy, Abriele Qudsi, Damian Ziaber, Earl Barr, and Sergey Mechtaev · 2024
Closest in time.
Repohyper: Better context retrieval is all you need for repository-level code completion
Huy N. Phan, Hoang N. Phan, Tien N. Nguyen, and Nghi D. Q. Bui · 2024
Closest in time.
Coverup: Coverage-guided llm-based test generation
Juan Altmayer Pizzorno and Emery D. Berger · 2024
Closest in time.
Syntactic robustness for llm-based code generation
Laboni Sarker, Mara Downing, Achintya Desai, and Tevfik Bultan · 2024
Closest in time.
Design2code: How far are we from automating front-end engineering?
Chenglei Si, Yanzhe Zhang, Zhengyuan Yang, Ruibo Liu, and Diyi Yang · 2024
Closest in time.
A comprehensive study of the capabilities of large language models for vulnerability detection
Benjamin Steenhoek, Md Mahbubur Rahman, Monoshi Kumar Roy, Mirza Sanjida Alam, Earl T. Barr, and Wei Le · 2024
Closest in time.
Llm4vuln: A unified evaluation framework for decoupling and enhancing llms’ vulnerability reasoning
Yuqiang Sun, Daoyuan Wu, Yue Xue, Han Liu, Wei Ma, Lyuye Zhang, Miaolei Shi, and Yang Liu · 2024
Closest in time.
Bugs in large language models generated code: An empirical study
Florian Tambon, Arghavan Moradi Dakhel, Amin Nikanjam, Foutse Khomh, Michel C. Desmarais, and Giuliano Antoniol · 2024
Closest in time.
Llm4decompile: Decompiling binary code with large language models
Hanzhuo Tan, Qi Luo, Jing Li, and Yuqun Zhang · 2024
Closest in time.
MAGIS: llm-based multi-agent framework for github issue resolution
Wei Tao, Yucheng Zhou, Wenqiang Zhang, and Yu Cheng · 2024
Closest in time.
Codegemma: Open code models based on gemma, 2024
CodeGemma Team · 2024
Closest in time.
Debugbench: Evaluating debugging capability of large language models
Runchu Tian, Yining Ye, Yujia Qin, Xin Cong, Yankai Lin, Yinxu Pan, Yesai Wu, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Tim van Dam, Frank van der Heijden, Philippe de Bekker, Berend Nieuwschepen, Marc Otten, and Maliheh Izadi · 2024
Closest in time.
Does your neural code completion model use my code? a membership inference approach
Yao Wan, Guanghua Wan, Shijie Zhang, Hongyu Zhang, Yulei Sui, Pan Zhou, Hai Jin, and Lichao Sun · 2024
Closest in time.
Benchmarking the communication competence of code generation for llms and llm agent
Jie JW Wu and Fatemeh H Fard · 2024
Closest in time.
Rui Xie, Zhengran Zeng, Zhuohao Yu, Chang Gao, Shikun Zhang, and Wei Ye · 2024
Closest in time.
WizardLM: Empowering large pre-trained language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang · 2024
Closest in time.
Yi: Open foundation models by 01.ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai · 2024
Closest in time.
Shifting the lens: Detecting malware in npm ecosystem with large language models
Nusrat Zahan, Philipp Burckhardt, Mikola Lysenko, Feross Aboukhadijeh, and Laurie A. Williams · 2024
Closest in time.
Codes: Natural language to code repository via multi-layer sketch
Daoguang Zan, Ailun Yu, Wei Liu, Dong Chen, Bo Shen, Wei Li, Yafen Yao, Yongshun Gong, Xiaolin Chen, Bei Guan, Zhiguang Yang, Yongji Wang, Qianxiang Wang, and Lizhen Cui · 2024
Closest in time.
Opencodeinterpreter: Integrating code generation with execution and refinement
Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue · 2024
Closest in time.
Large language model for vulnerability detection and repair: Literature review and the road ahead
Xin Zhou, Sicong Cao, Xiaobing Sun, and David Lo · 2024
Closest in time.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang, Peiyi Wang, Runxin Xu, Y Wu, Yukun Li, Huazuo Gao, Shirong Ma, et al · 2024
Closest in time.
Astraios: Parameter-efficient instruction tuning code large language models
Terry Yue Zhuo, Armel Zebaze, Nitchakarn Suppattarachai, Leandro von Werra, Harm de Vries, Qian Liu, and Niklas Muennighoff · 2024
Closest in time.
A neural architecture for generating natural language descriptions from source code changes
Pablo Loyola, Edison Marrese-Taylor, and Yutaka Matsuo · 2045
Closest in time.
A parallel corpus of python functions and documentation strings for automated code documentation and code generation
Antonio Valerio Miceli Barone and Rico Sennrich · 2053
Closest in time.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2074
Closest in time.
A convolutional attention network for extreme summarization of source code
Miltiadis Allamanis, Hao Peng, and Charles Sutton · 2091
Closest in time.
Typesql: Knowledge-based type-aware neural text-to-sql generation
Tao Yu, Zifan Li, Zilin Zhang, Rui Zhang, and Dragomir R. Radev · 2093
Closest in time.