Fetching the paper…
Reading the bibliography…
Statistical language modeling techniques have successfully been applied to large source code corpora, yielding a variety of new software development tools, such as tools for code suggestion, improving readability, and API migration.
Maybe Deep Neural Networks are the Best Choice for Modeling Source Code
Rafael-Michael Karampatsis and Charles A. Sutton. 2019 · 1903
Earlier work this paper cites.
Modeling Vocabulary for Big Code Machine Learning
Hlib Babii, Andrea Janes, and Romain Robbes. 2019 · 1904
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
A New Algorithm for Data Compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F Chen and Joshua Goodman. 1999 · 1999
Earlier work this paper cites.
Modelling Out-of-vocabulary Words for Robust Speech Recognition
Issam Bazzi. 2002 · 2002
Earlier work this paper cites.
A Neural Probabilistic Language Model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
The Porter stemming algorithm: then and now
Peter Willett. 2006 · 2006
Earlier work this paper cites.
Morph-based speech recognition and modeling of out-of-vocabulary words across languages
Mathias Creutz, Teemu Hirsimäki, Mikko Kurimo, Antti Puurula, Janne Pylkkönen, Vesa Siivola, Matti Varjokallio, Ebru Arisoy, Murat Saraçlar, and Andreas Stolcke. 2007 · 2007
Earlier work this paper cites.
To CamelCase or Under_score. In Proceedings of ICPC 2009 . 158–167
David W. Binkley, Marcia Davis, Dawn J. Lawrie, and Christopher Morrell. 2009 · 2009
Earlier work this paper cites.
Learning from examples to improve code completion systems. In Proceedings of ESEC/FSE 2009 . 213–222
Marcel Bruch, Martin Monperrus, and Mira Mezini. 2009 · 2009
Earlier work this paper cites.
Mining source code to automatically split identifiers for software analysis. In Proceedings of MSR 2009 . 71–80
Eric Enslen, Emily Hill, Lori L. Pollock, and K. Vijay-Shanker. 2009 · 2009
Earlier work this paper cites.
A study of the uniqueness of source code. In Proceedings of SIGSOFT/FSE 2010 . 147–156
Mark Gabel and Zhendong Su. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model. In Proceedings of INTERSPEECH 2010 . 1045–1048
Tomas Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
LINSEN: An efficient approach to split identifiers and expand abbreviations. In Proceedings of ICSM 2012 . 233–242
Anna Corazza, Sergio Di Martino, and Valerio Maggio. 2012 · 2012
Earlier work this paper cites.
On the Naturalness of Software. In Proceedings of ICSE 2012 . 837–847
Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. 2012 · 2012
Earlier work this paper cites.
Subword Language Modeling With Neural Networks
Tomas Mikolov, Ilya Sutskever, Anoop Deoras, Le Hai Son, Stefan Kombrink, and Jan Cernock. 2012 · 2012
Earlier work this paper cites.
Mining source code repositories at massive scale using language modeling. In Proceedings of MSR 2013 . 207–216
Miltiadis Allamanis and Charles A. Sutton. 2013 · 2013
Earlier work this paper cites.
Statistical NLP for computer program source code: An information theoretic perspective on programming language verbosity
Sergey Dudoladov. 2013 · 2013
Earlier work this paper cites.
Better Word Representations with Recursive Neural Networks for Morphology. In Proceedings of CoNLL 2013 . 104–113
Thang Luong, Richard Socher, and Christopher D. Manning. 2013 · 2013
Earlier work this paper cites.
Distributed Representations of Words and Phrases and Their Compositionality. In Proceedings of NIPS 2013 . USA, 3111–3119
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Lexical Statistical Machine Translation for Language Migration. In Proceedings ESEC/FSE 2013 . 651–654
Anh Tuan Nguyen, Tung Thanh Nguyen, and Tien N. Nguyen. 2013a · 2013
Earlier work this paper cites.
A Statistical Semantic Language Model for Source Code. In Proceedings of ESEC/FSE 2013 . New York, NY, USA, 532–542
Tung Thanh Nguyen, Anh Tuan Nguyen, Hoan Anh Nguyen, and Tien N. Nguyen. 2013b · 2013
Earlier work this paper cites.
Learning natural coding conventions. In Proceedings of SIGSOFT/FSE 2014 . 281–293
Miltiadis Allamanis, Earl T. Barr, Christian Bird, and Charles A. Sutton. 2014 · 2014
Earlier work this paper cites.
Syntax Errors Just Aren’t Natural: Improving Error Reporting with Language Models. In Proceedings of MSR 2014 . 252–261
Joshua Charles Campbell, Abram Hindle, and José Nelson Amaral. 2014 · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In Proceedings of EMNLP 2014 . 1724–1734
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
Junyoung Chung, Çaglar Gülçehre, KyungHyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
An empirical study of identifier splitting techniques
Emily Hill, David Binkley, Dawn Lawrie, Lori Pollock, and K Vijay-Shanker. 2014 · 2014
Earlier work this paper cites.
Defects4J: A database of existing faults to enable controlled testing studies for Java programs. In Proceedings of ISSTA 2014 . 437–440
René Just, Darioush Jalali, and Michael D Ernst. 2014 · 2014
Earlier work this paper cites.
Phrase-Based Statistical Translation of Programming Languages. In Proceedings of Onward! 2014 . 173–184
Svetoslav Karaivanov, Veselin Raychev, and Martin Vechev. 2014 · 2014
Cited alongside, same era.
Code Completion with Statistical Language Models. In Proceedings of PLDI 2014 . 419–428
Veselin Raychev, Martin Vechev, and Eran Yahav. 2014 · 2014
Cited alongside, same era.
On the localness of software. In Proceedings of SIGSOFT/FSE 2014 . 269–280
Zhaopeng Tu, Zhendong Su, and Premkumar T. Devanbu. 2014 · 2014
Cited alongside, same era.
Suggesting accurate method and class names. In Proceedings of ESEC/FSE 2015 . 38–49
Miltiadis Allamanis, Earl T. Barr, Christian Bird, and Charles A. Sutton. 2015a · 2015
Cited alongside, same era.
Bimodal Modelling of Source Code and Natural Language. In Proceedings of ICML 2015 , Vol. 37. 2123–2132
Miltiadis Allamanis, Daniel Tarlow, Andrew D. Gordon, and Yi Wei. 2015b · 2015
Cited alongside, same era.
Recovering clear, natural identifiers from obfuscated JS names. In Proceedings of ESEC/FSE 2017 . 683–693
Bogdan Vasilescu, Casey Casalnuovo, and Premkumar Devanbu. 2017 · 2017
Later among the works it cites.
Attention is all you need. In Proceedings of NIPS 2017 . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
A Survey of Machine Learning for Big Code and Naturalness
Miltiadis Allamanis, Earl T. Barr, Premkumar T. Devanbu, and Charles A. Sutton. 2018 · 2018
Later among the works it cites.
Context2Name: A Deep Learning-Based Approach to Infer Natural Variable Names from Usage Contexts
Rohan Bavishi, Michael Pradel, and Koushik Sen. 2018 · 2018
Later among the works it cites.
Sequencer: Sequence-to-sequence learning for end-to-end program repair
Zimin Chen, Steve Kommrusch, Michele Tufano, Louis-Noël Pouchet, Denys Poshyvanyk, and Martin Monperrus. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An Investigation of Statistical Language Modelling of Different Programming Language Types Using Large Corpora
Stefan Fiott. 2015 · 2015
Cited alongside, same era.
CACHECA: A Cache Language Model Based Code Suggestion Tool. In Proceedings of ICSE 2015 (Volume 2) . 705–708
Christine Franks, Zhaopeng Tu, Premkumar T. Devanbu, and Vincent Hellendoorn. 2015 · 2015
Cited alongside, same era.
On Using Very Large Target Vocabulary for Neural Machine Translation. In Proceedings of ACL 2015 . 1–10
Sébastien Jean, KyungHyun Cho, Roland Memisevic, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Pointer Networks. In Proceedings of NIPS 2015 . 2692–2700
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015 · 2015
Cited alongside, same era.
Toward Deep Learning Software Repositories. In Proceedings MSR 2015 . 334–345
Martin White, Christopher Vendome, Mario Linares-Vásquez, and Denys Poshyvanyk. 2015 · 2015
Cited alongside, same era.
Automated Correction for Syntax Errors in Programming Assignments using Recurrent Neural Networks
Sahil Bhatia and Rishabh Singh. 2016 · 2016
Cited alongside, same era.
PHOG: Probabilistic Model for Code. In Proceedings of ICML 2016 , Vol. 48. 2933–2942
Pavol Bielik, Veselin Raychev, and Martin T. Vechev. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
FRAGE: Frequency-Agnostic Word Representation. In Proceedings of NeurIPS 2018 . 1341–1352
ChengYue Gong, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2018 · 2018
Later among the works it cites.
Deep Code Comment Generation. In Proceedings of ICPC 2018 . 200–210
Xing Hu, Ge Li, Xin Xia, David Lo, and Zhi Jin. 2018 · 2018
Later among the works it cites.
Spiral: splitters for identifiers in source code files
Michael Hucka. 2018 · 2018
Later among the works it cites.
Meaningful Variable Names for Decompiled Code: A Machine Translation Approach. In Proceedings of ICPC 2018 . 20–30
Alan Jaffe, Jeremy Lacomis, Edward J Schwartz, Claire Le Goues, and Bogdan Vasilescu. 2018 · 2018
Later among the works it cites.
Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context. In Proceedings of ACL 2018 . 284–294
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Later among the works it cites.
Code Completion with Neural Attention and Pointer Networks. In Proceedings of IJCAI 2018 . 4159–4165
Jian Li, Yue Wang, Michael R. Lyu, and Irwin King. 2018 · 2018
Later among the works it cites.
Splitting source code identifiers using Bidirectional LSTM Recurrent Neural Network
Vadim Markovtsev, Waren Long, Egor Bulychev, Romain Keramitas, Konstantin Slavnov, and Gabor Markowski. 2018 · 2018
Later among the works it cites.
Deep Contextualized Word Representations. In Proceedings of NAACL-HLT 2018 . 2227–2237
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
DeepBugs: A Learning Approach to Name-based Bug Detection
Michael Pradel and Koushik Sen. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Syntax and Sensibility: Using language models to detect and correct syntax errors. In Proceedings of SANER 2018 . 311–322
Eddie Antonio Santos, Joshua Charles Campbell, Dhvani Patel, Abram Hindle, and José Nelson Amaral. 2018 · 2018
Later among the works it cites.
Learning to mine aligned code and natural language pairs from stack overflow. In Proceedings of MSR 2018 . 476–486
Pengcheng Yin, Bowen Deng, Edgar Chen, Bogdan Vasilescu, and Graham Neubig. 2018 · 2018
Later among the works it cites.
The adverse effects of code duplication in machine learning models of code. In Proceedings of Onward! 2019 . 143–153
Miltiadis Allamanis. 2019 · 2019
Later among the works it cites.
code2seq: Generating Sequences from Structured Representations of Code. In Proceedings of ICLR 2019
Uri Alon, Shaked Brody, Omer Levy, and Eran Yahav. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT 2019 . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
When code completion fails: a case study on real-world completions. In Proceedings of ICSE 2019 . 960–970
Vincent J. Hellendoorn, Sebastian Proksch, Harald C. Gall, and Alberto Bacchelli. 2019 · 2019
Later among the works it cites.
Universal language model fine-tuning for text classification. In Proceedings of ACL 2019 . 328–339
Jeremy Howard and Sebastian Ruder. 2018 · 2019
Later among the works it cites.
NL2Type: inferring JavaScript function types from natural language information. In Proceedings of ICSE 2019 . 304–315
Rabee Sohail Malik, Jibesh Patra, and Michael Pradel. 2019 · 2019
Later among the works it cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Leo Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Natural software revisited. In Proceedings of ICSE 2019 . 37–48
Musfiqur Rahman, Dharani Palani, and Peter C. Rigby. 2019 · 2019
Later among the works it cites.
Leveraging small software engineering data sets with pre-trained neural networks. In Proceedings of ICSE (NIER) 2019 . 29–32
Romain Robbes and Andrea Janes. 2019 · 2019
Later among the works it cites.
On Learning Meaningful Code Changes via Neural Machine Translation. In Proceedings of ICSE 2019 . 25–36
Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk. 2019 · 2019
Later among the works it cites.
From Characters to Words to in Between: Do We Capture Morphology?. In Proceedings of ACL 2017 . 2016–2027
Clara Vania and Adam Lopez. 2017 · 2027
Closest in time.
Summarizing Source Code using a Neural Attention Model. In Proceedings of ACL 2016 . 2073–2083
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016 · 2083
Closest in time.
A Convolutional Attention Network for Extreme Summarization of Source Code. In Proceedings of ICML 2016 , Vol. 48. 2091–2100
Miltiadis Allamanis, Hao Peng, and Charles A. Sutton. 2016 · 2091
Closest in time.