Fetching the paper…
Reading the bibliography…
Benchmark datasets have a significant impact on accelerating research in programming language tasks.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Populating a release history database from version control and bug tracking systems. In International Conference on Software Maintenance, 2003. ICSM 2003. Proceedings. IEEE, 23–32
Michael Fischer, Martin Pinzger, and Harald Gall. 2003 · 2003
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz J Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
Locality-sensitive hashing scheme based on p-stable distributions. In Proceedings of the twentieth annual symposium on Computational geometry . 253–262
Mayur Datar, Nicole Immorlica, Piotr Indyk, and Vahab S Mirrokni. 2004 · 2004
Earlier work this paper cites.
Orange: a method for evaluating automatic evaluation metrics for machine translation. In COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics . 501–507
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Deckard: Scalable and accurate tree-based detection of code clones. In 29th International Conference on Software Engineering (ICSE’07) . IEEE, 96–105
Lingxiao Jiang, Ghassan Misherghi, Zhendong Su, and Stephane Glondu. 2007 · 2007
Earlier work this paper cites.
Learning from Examples to Improve Code Completion Systems. In Proceedings of the 7th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on The Foundations of Software Engineering (Amsterdam, The Netherlands) (ESEC/FSE ’09) . Association for Computing Machinery, New York, NY, USA, 213–222
Marcel Bruch, Martin Monperrus, and Mira Mezini. 2009 · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition . Ieee, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Evosuite: automatic test suite generation for object-oriented software. In Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering . 416–419
Gordon Fraser and Andrea Arcuri. 2011 · 2011
Earlier work this paper cites.
On the naturalness of software. In 2012 34th International Conference on Software Engineering (ICSE) . IEEE, 837–847
Abram Hindle, Earl T Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. 2012 · 2012
Earlier work this paper cites.
Mining Source Code Repositories at Massive Scale using Language Modeling. In 2013 10th Working Conference on Mining Software Repositories (MSR) . IEEE, 207–216
Miltiadis Allamanis and Charles Sutton. 2013 · 2013
Earlier work this paper cites.
Mining idioms from source code. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering . 472–483
Miltiadis Allamanis and Charles Sutton. 2014 · 2014
Earlier work this paper cites.
Fine-grained and accurate source code differencing. In Proceedings of the 29th ACM/IEEE international conference on Automated software engineering . 313–324
Jean-Rémy Falleri, Floréal Morandat, Xavier Blanc, Matias Martinez, and Martin Monperrus. 2014 · 2014
Earlier work this paper cites.
Phrase-Based Statistical Translation of Programming Languages. In Proceedings of the 2014 ACM International Symposium on New Ideas, New Paradigms, and Reflections on Programming Software (Portland, Oregon, USA) (Onward! 2014) . Association for Computing Machinery, New York, NY, USA, 173–184
Svetoslav Karaivanov, Veselin Raychev, and Martin Vechev. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Code Completion with Statistical Language Models. In Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation (Edinburgh, United Kingdom) (PLDI ’14) . Association for Computing Machinery, New York, NY, USA, 419–428
Veselin Raychev, Martin Vechev, and Eran Yahav. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks. In Advances in neural information processing systems . 3104–3112
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Towards a big data curated benchmark of inter-project code clones. In 2014 IEEE International Conference on Software Maintenance and Evolution . IEEE, 476–480
Jeffrey Svajlenko, Judith F Islam, Iman Keivanloo, Chanchal K Roy, and Mohammad Mamun Mia. 2014 · 2014
Earlier work this paper cites.
Synthesizing Data Structure Transformations from Input-Output Examples. In Proceedings of the 36th ACM SIGPLAN Conference on Programming Language Design and Implementation (Portland, OR, USA) (PLDI ’15) . Association for Computing Machinery, New York, NY, USA, 229–239
John K. Feser, Swarat Chaudhuri, and Isil Dillig. 2015 · 2015
Earlier work this paper cites.
Neural programmer: Inducing latent programs with gradient descent
Arvind Neelakantan, Quoc V Le, and Ilya Sutskever. 2015 · 2015
Earlier work this paper cites.
Divide-and-conquer approach for multi-phase statistical migration for source code (t). In 2015 30th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 585–596
Anh Tuan Nguyen, Tung Thanh Nguyen, and Tien N Nguyen. 2015 · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5206–5210
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Neural programmer-interpreters
Scott Reed and Nando De Freitas. 2015 · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Earlier work this paper cites.
Predicting a correct program in programming by example. In International Conference on Computer Aided Verification . Springer, 398–414
Rishabh Singh and Sumit Gulwani. 2015 · 2015
Earlier work this paper cites.
PHOG: Probabilistic Model for Code. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 (New York, NY, USA) (ICML’16) . JMLR.org, 2933–2942
Pavol Bielik, Veselin Raychev, and Martin Vechev. 2016 · 2016
Earlier work this paper cites.
Convolutional neural networks over tree structures for programming language processing. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence . 1287–1293
Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
On the" naturalness" of buggy code. In 2016 IEEE/ACM 38th International Conference on Software Engineering (ICSE) . IEEE, 428–439
Baishakhi Ray, Vincent Hellendoorn, Saheel Godhane, Zhaopeng Tu, Alberto Bacchelli, and Premkumar Devanbu. 2016 · 2016
Earlier work this paper cites.
Probabilistic Model for Code with Decision Trees
Veselin Raychev, Pavol Bielik, and Martin Vechev. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1715–1725
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Refinement Types for TypeScript. In Proceedings of the 37th ACM SIGPLAN Conference on Programming Language Design and Implementation (Santa Barbara, CA, USA) (PLDI ’16) . Association for Computing Machinery, New York, NY, USA, 310–325
Panagiotis Vekris, Benjamin Cosman, and Ranjit Jhala. 2016 · 2016
Cited alongside, same era.
Deep learning code fragments for code clone detection. In 2016 31st IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 87–98
Martin White, Michele Tufano, Christopher Vendome, and Denys Poshyvanyk. 2016 · 2016
Cited alongside, same era.
Learning to represent programs with graphs
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi. 2017 · 2017
Cited alongside, same era.
Program synthesis
Sumit Gulwani, Oleksandr Polozov, Rishabh Singh, et al · 2017
Cited alongside, same era.
Spoc: Search-based pseudocode to code. In Advances in Neural Information Processing Systems . 11906–11917
Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, and Percy S Liang. 2019 · 2019
Later among the works it cites.
Improving bug detection via context-based code representation learning and attention-based neural networks
Yi Li, Shaohua Wang, Tien N Nguyen, and Son Van Nguyen. 2019 · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Later among the works it cites.
Aroma: Code recommendation via structural code search
Sifei Luan, Di Yang, Celeste Barnaby, Koushik Sen, and Satish Chandra. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rahul Gupta, Soham Pal, Aditya Kanade, and Shirish Shevade. 2017 · 2017
Cited alongside, same era.
Are deep neural networks the best choice for modeling source code?. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering . 763–773
Vincent J Hellendoorn and Premkumar Devanbu. 2017 · 2017
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al · 2017
Cited alongside, same era.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bayesian Sketch Learning for Program Synthesis
Murali Vijayaraghavan, Chaudhuri Swarat, and Jermaine Chris. 2017 · 2017
Cited alongside, same era.
Supervised Deep Features for Software Functional Clone Detection by Exploiting Lexical and Syntactical Information in Source Code.. In IJCAI . 3034–3040
Huihui Wei and Ming Li. 2017 · 2017
Cited alongside, same era.
A syntactic neural model for general-purpose code generation
Pengcheng Yin and Graham Neubig. 2017 · 2017
Cited alongside, same era.
A Survey of Machine Learning for Big Code and Naturalness
Miltiadis Allamanis, Earl T. Barr, Premkumar Devanbu, and Charles Sutton. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al · 2019
Later among the works it cites.
Pythia: ai-assisted code completion system. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 2727–2735
Alexey Svyatkovskiy, Ying Zhao, Shengyu Fu, and Neel Sundaresan. 2019 · 2019
Later among the works it cites.
An empirical study on learning bug-fixing patches in the wild via neural machine translation
Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. 2019 · 2019
Later among the works it cites.
Neural program repair by jointly learning to localize and repair
Marko Vasic, Aditya Kanade, Petros Maniatis, David Bieber, and Rishabh Singh. 2019 · 2019
Later among the works it cites.
Code generation as a dual task of code summarization. In Advances in Neural Information Processing Systems . 6563–6573
Bolin Wei, Ge Li, Xin Xia, Zhiyi Fu, and Zhi Jin. 2019 · 2019
Later among the works it cites.
Neural detection of semantic code clones via tree-based convolution. In 2019 IEEE/ACM 27th International Conference on Program Comprehension (ICPC) . IEEE, 70–80
Hao Yu, Wing Lam, Long Chen, Ge Li, Tao Xie, and Qianxiang Wang. 2019 · 2019
Later among the works it cites.
A novel neural source code representation based on abstract syntax tree. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE) . IEEE, 783–794
Jian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun, Kaixuan Wang, and Xudong Liu. 2019 · 2019
Later among the works it cites.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks. In Advances in Neural Information Processing Systems . 10197–10207
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu. 2019 · 2019
Later among the works it cites.
PyMT5: multi-mode translation of natural language and Python code with transformers
Colin B Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan. 2020 · 2020
Later among the works it cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Later among the works it cites.
GraphCodeBERT: Pre-training Code Representations with Data Flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Jian Yin, Daxin Jiang, et al · 2020
Later among the works it cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Later among the works it cites.
Big Code!= Big Vocabulary: Open-Vocabulary Models for Source Code
Rafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton, and Andrea Janes. 2020 · 2020
Later among the works it cites.
Unsupervised Translation of Programming Languages
Marie-Anne Lachaux, Baptiste Roziere, Lowik Chanussot, and Guillaume Lample. 2020 · 2020
Later among the works it cites.
Xglue: A new benchmark dataset for cross-lingual pre-training, understanding and generation
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, et al · 2020
Later among the works it cites.
Semantic Code Search via Equational Reasoning. In Proceedings of the 41st ACM SIGPLAN Conference on Programming Language Design and Implementation (London, UK) (PLDI 2020) . Association for Computing Machinery, New York, NY, USA, 1066–1082
Varot Premtoon, James Koppel, and Armando Solar-Lezama. 2020 · 2020
Later among the works it cites.
CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2020
Later among the works it cites.
IntelliCode Compose: Code Generation Using Transformer
Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020 · 2020
Later among the works it cites.
Unit Test Case Generation with Transformers
Michele Tufano, Dawn Drain, Alexey Svyatkovskiy, Shao Kun Deng, and Neel Sundaresan. 2020 · 2020
Later among the works it cites.
Detecting Code Clones with Graph Neural Network and Flow-Augmented Abstract Syntax Tree. In 2020 IEEE 27th International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 261–271
Wenhan Wang, Ge Li, Bo Ma, Xin Xia, and Zhi Jin. 2020b · 2020
Later among the works it cites.
TranSˆ 3: A Transformer-based Framework for Unifying Code Summarization and Code Search
Wenhua Wang, Yuqun Zhang, Zhengran Zeng, and Guandong Xu. 2020c · 2020
Later among the works it cites.
Incorporating external knowledge through pre-training for natural language to code generation
Frank F Xu, Zhengbao Jiang, Pengcheng Yin, Bogdan Vasilescu, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Are the Code Snippets What We Are Searching for? A Benchmark and an Empirical Study on Code Search with Natural-Language Queries. In 2020 IEEE 27th International Conference on Software Analysis, Evolution and Reengineering (SANER) . 344–354
S. Yan, H. Yu, Y. Chen, B. Shen, and L. Jiang. 2020 · 2020
Later among the works it cites.
MISIM: An End-to-End Neural Code Similarity System
Fangke Ye, Shengtian Zhou, Anand Venkat, Ryan Marucs, Nesime Tatbul, Jesmin Jahan Tithi, Paul Petersen, Timothy Mattson, Tim Kraska, Pradeep Dubey, et al · 2020
Later among the works it cites.
Semantic Scaffolds for Pseudocode-to-Code Generation
Ruiqi Zhong, Mitchell Stern, and Dan Klein. 2020 · 2020
Later among the works it cites.
Summarizing source code using a neural attention model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2073–2083
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016 · 2083
Closest in time.
A convolutional attention network for extreme summarization of source code. In International conference on machine learning . 2091–2100
Miltiadis Allamanis, Hao Peng, and Charles Sutton. 2016 · 2091
Closest in time.