Fetching the paper…
Reading the bibliography…
We present a neural model for representing snippets of code as continuous distributed vectors ("code embeddings").
Distributional structure
Zellig S Harris. 1954 · 1954
Earlier work this paper cites.
A Synopsis of Linguistic Theory, 1930-1955
J.R. Firth. 1957 · 1955
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
A vector space model for automatic indexing
Gerard Salton, Anita Wong, and Chung-Shu Yang. 1975 · 1975
Earlier work this paper cites.
Indexing by latent semantic analysis
Scott Deerwester, Susan T Dumais, George W Furnas, Thomas K Landauer, and Richard Harshman. 1990 · 1990
Earlier work this paper cites.
Refactoring: improving the design of existing code
Martin Fowler and Kent Beck. 1999 · 1999
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Reuven Rubinstein. 1999 · 1999
Earlier work this paper cites.
Combinatorial optimization, cross-entropy, ants and rare events
Reuven Y Rubinstein. 2001 · 2001
Earlier work this paper cites.
A Neural Probabilistic Language Model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
Re-evaluation the role of bleu in machine translation research. In 11th Conference of the European Chapter of the Association for Computational Linguistics
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Multi-label classification: An overview
Grigorios Tsoumakas and Ioannis Katakis. 2006 · 2006
Earlier work this paper cites.
Similarity of semantic relations
Peter D Turney. 2006 · 2006
Earlier work this paper cites.
A Unified Architecture for Natural Language Processing: Deep Neural Networks with Multitask Learning. In Proceedings of the 25th International Conference on Machine Learning (ICML ’08) . ACM, New York, NY, USA, 160–167
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
Debugging Method Names. In Proceedings of the 23rd European Conference on ECOOP 2009 — Object-Oriented Programming (Genoa) . Springer-Verlag, Berlin, Heidelberg, 294–317
Einar W. Høst and Bjarte M. Østvold. 2009 · 2009
Earlier work this paper cites.
Neural conditional random fields. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics . 177–184
Thierry Artieres et al · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics . 249–256
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Word Representations: A Simple and General Method for Semi-supervised Learning. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics (ACL ’10) . Association for Computational Linguistics, Stroudsburg, PA, USA, 384–394
Joseph Turian, Lev Ratinov, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Domain adaptation for large-scale sentiment classification: A deep learning approach. In Proceedings of the 28th international conference on machine learning (ICML-11) . 513–520
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011 · 2011
Earlier work this paper cites.
Parsing Natural Scenes and Natural Language with Recursive Neural Networks. In Proceedings of the 26th International Conference on Machine Learning (ICML)
Richard Socher, Cliff C. Lin, Andrew Y. Ng, and Christopher D. Manning. 2011 · 2011
Earlier work this paper cites.
On the Naturalness of Software. In Proceedings of the 34th International Conference on Software Engineering (ICSE ’12) . IEEE Press, Piscataway, NJ, USA, 837–847
Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. 2012 · 2012
Earlier work this paper cites.
Typestate-based Semantic Code Search over Partial Programs. In Proceedings of the ACM International Conference on Object Oriented Programming Systems Languages and Applications (OOPSLA ’12) . ACM, New York, NY, USA, 997–1016
Alon Mishne, Sharon Shoham, and Eran Yahav. 2012 · 2012
Earlier work this paper cites.
Mining Source Code Repositories at Massive Scale Using Language Modeling. In Proceedings of the 10th Working Conference on Mining Software Repositories (MSR ’13) . IEEE Press, Piscataway, NJ, USA, 207–216
Miltiadis Allamanis and Charles Sutton. 2013 · 2013
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a · 2013
Earlier work this paper cites.
Natural language models for predicting programming comments
Dana Movshovitz-Attias and William W Cohen. 2013 · 2013
Cited alongside, same era.
A Statistical Semantic Language Model for Source Code. In Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering (ESEC/FSE 2013) . ACM, New York, NY, USA, 532–542
Tung Thanh Nguyen, Anh Tuan Nguyen, Hoan Anh Nguyen, and Tien N. Nguyen. 2013 · 2013
Cited alongside, same era.
Learning Natural Coding Conventions. In Proceedings of the 22Nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE 2014) . ACM, New York, NY, USA, 281–293
Miltiadis Allamanis, Earl T. Barr, Christian Bird, and Charles Sutton. 2014 · 2014
Cited alongside, same era.
Mining Idioms from Source Code. In Proceedings of the 22Nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE 2014) . ACM, New York, NY, USA, 472–483
Miltiadis Allamanis and Charles Sutton. 2014 · 2014
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition. In Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on . IEEE, 4945–4949
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio. 2016 · 2016
Later among the works it cites.
PHOG: Probabilistic Model for Code. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016 . 2933–2942
Pavol Bielik, Veselin Raychev, and Martin T. Vechev. 2016 · 2016
Later among the works it cites.
Statistical Similarity in Binaries. In PLDI’16: Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation
Yaniv David, Nimrod Partush, and Eran Yahav. 2016 · 2016
Later among the works it cites.
Summarizing Source Code using a Neural Attention Model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jimmy Ba, Volodymyr Mnih, and Koray Kavukcuoglu. 2014 · 2014
Cited alongside, same era.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Tracelet-Based Code Search in Executables. In PLDI’14: Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation . 349–360
Yaniv David and Eran Yahav. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Distributed Representations of Sentences and Documents. In Proceedings of the 31st International Conference on Machine Learning (ICML-14) , Tony Jebara and Eric P. Xing (Eds.). JMLR Workshop and Conference Proceedings, 1188–1196
Quoc Le and Tomas Mikolov. 2014 · 2014
Cited alongside, same era.
Neural Word Embeddings as Implicit Matrix Factorization. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada . 2177–2185
Omer Levy and Yoav Goldberg. 2014b · 2014
Cited alongside, same era.
Structured Generative Models of Natural Source Code. In Proceedings of the 31st International Conference on International Conference on Machine Learning - Volume 32 (ICML’14) . JMLR.org, II–649–II–657
Chris J. Maddison and Daniel Tarlow. 2014 · 2014
Cited alongside, same era.
Recurrent Models of Visual Attention. In Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS’14) . MIT Press, Cambridge, MA, USA, 2204–2212
Volodymyr Mnih, Nicolas Heess, Alex Graves, and Koray Kavukcuoglu. 2014 · 2014
Cited alongside, same era.
Estimating Types in Executables using Predictive Modeling. In POPL’16: Proceedings of the ACM SIGPLAN Conference on Principles of Programming Languages
Omer Katz, Ran El-Yaniv, and Eran Yahav. 2016 · 2016
Later among the works it cites.
Probabilistic Model for Code with Decision Trees. In Proceedings of the 2016 ACM SIGPLAN International Conference on Object-Oriented Programming, Systems, Languages, and Applications (OOPSLA 2016) . ACM, New York, NY, USA, 731–747
Veselin Raychev, Pavol Bielik, and Martin Vechev. 2016 · 2016
Later among the works it cites.
Bidirectional attention flow for machine comprehension
Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2016 · 2016
Later among the works it cites.
Programming with "Big Code"
Martin T. Vechev and Eran Yahav. 2016 · 2016
Later among the works it cites.
Leveraging a Corpus of Natural Language Descriptions for Program Similarity. In Proceedings of the 2016 ACM International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software (Onward! 2016) . ACM, New York, NY, USA, 197–211
Meital Zilberstein and Eran Yahav. 2016 · 2016
Later among the works it cites.
A Survey of Machine Learning for Big Code and Naturalness
Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton. 2017 · 2017
Later among the works it cites.
Neural Attribute Machines for Program Generation
Matthew Amodio, Swarat Chaudhuri, and Thomas W. Reps. 2017 · 2017
Later among the works it cites.
Similarity of Binaries through re-optimization. In PLDI’17: Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation
Yaniv David, Nimrod Partush, and Eran Yahav. 2017 · 2017
Later among the works it cites.
Zero-Shot Relation Extraction via Reading Comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), Vancouver, Canada, August 3-4, 2017 . 333–342
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017 · 2017
Later among the works it cites.
DéJàVu: A Map of Code Duplicates on GitHub
Cristina V. Lopes, Petr Maj, Pedro Martins, Vaibhav Saini, Di Yang, Jakub Zitny, Hitesh Sajnani, and Jan Vitek. 2017 · 2017
Later among the works it cites.
Data-Driven Program Completion
Yanxin Lu, Swarat Chaudhuri, Chris Jermaine, and David Melski. 2017 · 2017
Later among the works it cites.
Bayesian Sketch Learning for Program Synthesis
Vijayaraghavan Murali, Swarat Chaudhuri, and Chris Jermaine. 2017 · 2017
Later among the works it cites.
Attention is all you need. In Advances in Neural Information Processing Systems . 6000–6010
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Learning to Represent Programs with Graphs. In ICLR
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi. 2018 · 2018
Closest in time.
A General Path-based Representation for Predicting Program Properties. In Proceedings of the 39th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI 2018) . ACM, New York, NY, USA, 404–419
Uri Alon, Meital Zilberstein, Omer Levy, and Eran Yahav. 2018 · 2018
Closest in time.
Statistical Reconstruction of Class Hierarchies in Binaries. In ASPLOS’18: Proceedings of the ACM Conference on Architectural Support for Programming Languages and Operating Systems
Omer Katz, Noam Rinetzky, and Eran Yahav. 2018 · 2018
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention. In International Conference on Machine Learning . 2048–2057
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.
A Convolutional Attention Network for Extreme Summarization of Source Code. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016 . 2091–2100
Miltiadis Allamanis, Hao Peng, and Charles A. Sutton. 2016 · 2091
Closest in time.