Fetching the paper…
Reading the bibliography…
When building machine learning models that operate on source code, several decisions have to be made to model source-code vocabulary.
Maybe Deep Neural Networks are the Best Choice for Modeling Source Code
Rafael-Michael Karampatsis and Charles Sutton. 2019 · 1903
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F Chen and Joshua Goodman. 1999 · 1999
Earlier work this paper cites.
Classes for fast maximum entropy training. In Acoustics, Speech, and Signal Processing, 2001. Proceedings.(ICASSP’01). 2001 IEEE International Conference on , Vol. 1. IEEE, 561–564
Joshua Goodman. 2001 · 2001
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003a · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model.. In Aistats , Vol. 5. Citeseer, 246–252
Frederic Morin and Yoshua Bengio. 2005 · 2005
Earlier work this paper cites.
Concise and consistent naming
Florian Deissenboeck and Markus Pizka. 2006 · 2006
Earlier work this paper cites.
The Porter stemming algorithm: then and now
Peter Willett. 2006 · 2006
Earlier work this paper cites.
Morph-based speech recognition and modeling of out-of-vocabulary words across languages
Mathias Creutz, Teemu Hirsimäki, Mikko Kurimo, Antti Puurula, Janne Pylkkönen, Vesa Siivola, Matti Varjokallio, Ebru Arisoy, Murat Saraçlar, and Andreas Stolcke. 2007 · 2007
Earlier work this paper cites.
To camelcase or under_score. In 2009 IEEE 17th International Conference on Program Comprehension (ICPC 2009) . IEEE, 158–167
Dave Binkley, Marcia Davis, Dawn Lawrie, and Christopher Morrell. 2009 · 2009
Earlier work this paper cites.
A study of the uniqueness of source code. In Proceedings of the eighteenth ACM SIGSOFT international symposium on Foundations of software engineering . ACM, 147–156
Mark Gabel and Zhendong Su. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model. In Eleventh annual conference of the international speech communication association
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
On the naturalness of software. In Software Engineering (ICSE), 2012 34th International Conference on . IEEE, 837–847
Abram Hindle, Earl T Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. 2012 · 2012
Earlier work this paper cites.
Subword language modeling with neural networks
Tomáš Mikolov, Ilya Sutskever, Anoop Deoras, Hai-Son Le, Stefan Kombrink, and Jan Cernocky. 2012 · 2012
Earlier work this paper cites.
LSTM neural networks for language modeling. In Thirteenth annual conference of the international speech communication association
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney. 2012 · 2012
Earlier work this paper cites.
Mining source code repositories at massive scale using language modeling. In Proceedings of the 10th Working Conference on Mining Software Repositories . IEEE Press, 207–216
Miltiadis Allamanis and Charles Sutton. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems . 3111–3119
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
A statistical semantic language model for source code. In Proceedings of the 9th Joint Meeting on Foundations of Software Engineering . ACM, 532–542
Tung Thanh Nguyen, Anh Tuan Nguyen, Hoan Anh Nguyen, and Tien N Nguyen. 2013 · 2013
Earlier work this paper cites.
Learning natural coding conventions. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering . ACM, 281–293
Miltiadis Allamanis, Earl T Barr, Christian Bird, and Charles Sutton. 2014 · 2014
Earlier work this paper cites.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Code completion with statistical language models. In Acm Sigplan Notices , Vol. 49. ACM, 419–428
Veselin Raychev, Martin Vechev, and Eran Yahav. 2014 · 2014
Earlier work this paper cites.
On the localness of software. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering . ACM, 269–280
Zhaopeng Tu, Zhendong Su, and Premkumar Devanbu. 2014 · 2014
Cited alongside, same era.
Suggesting accurate method and class names. In Proceedings of the 10th Joint Meeting on Foundations of Software Engineering . ACM, 38–49
Miltiadis Allamanis, Earl T Barr, Christian Bird, and Charles Sutton. 2015 · 2015
Cited alongside, same era.
Predicting program properties from big code. In ACM SIGPLAN Notices , Vol. 50. ACM, 111–124
Veselin Raychev, Martin Vechev, and Andreas Krause. 2015 · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Cited alongside, same era.
Toward deep learning software repositories. In Proceedings of the 12th Working Conference on Mining Software Repositories . IEEE Press, 334–345
A survey of machine learning for big code and naturalness
Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton. 2018a · 2018
Later among the works it cites.
code2seq: Generating sequences from structured representations of code
Uri Alon, Omer Levy, and Eran Yahav. 2018a · 2018
Later among the works it cites.
A general path-based representation for predicting program properties
Uri Alon, Meital Zilberstein, Omer Levy, and Eran Yahav. 2018b · 2018
Later among the works it cites.
Context2Name: A deep learning-based approach to infer natural variable names from usage contexts
Rohan Bavishi, Michael Pradel, and Koushik Sen. 2018 · 2018
Later among the works it cites.
Neural Code Comprehension: A Learnable Representation of Code Semantics
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Martin White, Christopher Vendome, Mario Linares-Vásquez, and Denys Poshyvanyk. 2015 · 2015
Cited alongside, same era.
Quasi-recurrent neural networks
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier. 2016 · 2016
Cited alongside, same era.
Deep API learning. In Proceedings of the 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering . ACM, 631–642
Xiaodong Gu, Hongyu Zhang, Dongmei Zhang, and Sunghun Kim. 2016 · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Cited alongside, same era.
Character-Aware Neural Language Models.. In AAAI . 2741–2749
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush. 2016 · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Unsupervised pretraining for sequence to sequence learning
Prajit Ramachandran, Peter J Liu, and Quoc V Le. 2016 · 2016
Cited alongside, same era.
Tal Ben-Nun, Alice Shoshana Jakobovits, and Torsten Hoefler. 2018 · 2018
Later among the works it cites.
Deep RNNs encode soft hierarchical syntax
Terra Blevins, Omer Levy, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Word embeddings for the software engineering domain. In Proceedings of the 15th International Conference on Mining Software Repositories . ACM, 38–41
Vasiliki Efstathiou, Christos Chatzilenas, and Diomidis Spinellis. 2018 · 2018
Later among the works it cites.
Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Vol. 1. 328–339
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Later among the works it cites.
Deep code comment generation. In Proceedings of the 26th Conference on Program Comprehension . ACM, 200–210
Xing Hu, Ge Li, Xin Xia, David Lo, and Zhi Jin. 2018 · 2018
Later among the works it cites.
Meaningful Variable Names for Decompiled Code: A Machine Translation Approach. In Proceedings of the 26th International Conference on Program Comprehension (ICPC). ACM
Alan Jaffe, Jeremy Lacomis, Edward J Schwartz, Claire Le Goues, and Bogdan Vasilescu. 2018 · 2018
Later among the works it cites.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Later among the works it cites.
A deep neural network language model with contexts for source code. In 2018 IEEE 25th International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 323–334
Anh Tuan Nguyen, Trong Duc Nguyen, Hung Dang Phan, and Tien N Nguyen. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Syntax and Sensibility: Using language models to detect and correct syntax errors. In 2018 IEEE 25th International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 311–322
Eddie Antonio Santos, Joshua Charles Campbell, Dhvani Patel, Abram Hindle, and José Nelson Amaral. 2018 · 2018
Later among the works it cites.
Deep Learning Similarities from Different Representations of Source Code. In Proceedings of the 15th International Conference on Mining Software Repositories
Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. 2018 · 2018
Later among the works it cites.
Learning to Mine Aligned Code and Natural Language Pairs from Stack Overflow
Pengcheng Yin, Bowen Deng, Edgar Chen, Bogdan Vasilescu, and Graham Neubig. 2018 · 2018
Later among the works it cites.
Code2Vec: Learning Distributed Representations of Code
Uri Alon, Meital Zilberstein, Omer Levy, and Eran Yahav. 2019 · 2019
Closest in time.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Leo Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Natural Software Revisited. In Proceedings of the 41st ACM/IEEE International Conference on Software Engineering . in press
Musfiqur Rahman, Dharani Kumar, and Peter Rigby. 2019 · 2019
Closest in time.
Leveraging Small Software Engineering Data Sets with Pre-trained Neural Networks. In Proceedings of the 41st ACM/IEEE International Conference on Software Engineering . in press
Romain Robbes and Andrea Janes. 2019 · 2019
Closest in time.
On Learning Meaningful Code Changes via Neural Machine Translation
Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk. 2019 · 2019
Closest in time.
A convolutional attention network for extreme summarization of source code. In Proceedings of the 33rd International Conference on Machine Learning . 2091–2100
Miltiadis Allamanis, Hao Peng, and Charles Sutton. 2016 · 2091
Closest in time.