Fetching the paper…
Reading the bibliography…
The field of big code relies on mining large corpora of code to perform some learning task.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
A survey on software clone detection research
Chanchal Kumar Roy and James R Cordy. 2007 · 2007
Earlier work this paper cites.
“Cloning considered harmful” considered harmful: patterns of cloning in software
Cory J Kapser and Michael W Godfrey. 2008 · 2008
Earlier work this paper cites.
code2seq: Generating Sequences from Structured Representations of Code. In Proceedings of the International Conference on Learning Representations (ICLR)
Uri Alon, Omer Levy, and Eran Yahav. 2010 · 2010
Earlier work this paper cites.
The Qualitas Corpus: A curated collection of Java code for empirical studies. In Software Engineering Conference (APSEC), 2010 17th Asia Pacific . IEEE, 336–345
Ewan Tempero, Craig Anslow, Jens Dietrich, Ted Han, Jing Li, Markus Lumpe, Hayden Melton, and James Noble. 2010 · 2010
Earlier work this paper cites.
On the naturalness of software. In Software Engineering (ICSE), 2012 34th International Conference on . IEEE, 837–847
Abram Hindle, Earl T Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. 2012 · 2012
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Kevin P Murphy. 2012 · 2012
Earlier work this paper cites.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton. 2012 · 2012
Earlier work this paper cites.
Mining source code repositories at massive scale using language modeling. In Proceedings of the Working Conference on Mining Software Repositories (MSR) . IEEE Press, 207–216
Miltiadis Allamanis and Charles Sutton. 2013 · 2013
Earlier work this paper cites.
Syntax errors just aren’t natural: improving error reporting with language models. In Proceedings of the Working Conference on Mining Software Repositories (MSR) . ACM, 252–261
Joshua Charles Campbell, Abram Hindle, and José Nelson Amaral. 2014 · 2014
Cited alongside, same era.
Structured generative models of natural source code. In Proceedings of the International Conference on Machine Learning (ICML) . 649–657
Chris Maddison and Daniel Tarlow. 2014 · 2014
Cited alongside, same era.
Code completion with statistical language models. In Proceedings of the Symposium on Programming Language Design and Implementation (PLDI) , Vol. 49. ACM, 419–428
Veselin Raychev, Martin Vechev, and Eran Yahav. 2014 · 2014
Cited alongside, same era.
Predicting program properties from Big Code. In Proceedings of the Symposium on Principles of Programming Languages (POPL) , Vol. 50. ACM, 111–124
Veselin Raychev, Martin Vechev, and Andreas Krause. 2015 · 2015
Cited alongside, same era.
DéjàVu: a map of code duplicates on GitHub
Cristina V Lopes, Petr Maj, Pedro Martins, Vaibhav Saini, Di Yang, Jakub Zitny, Hitesh Sajnani, and Jan Vitek. 2017 · 2017
Later among the works it cites.
A survey of machine learning for big code and naturalness
Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton. 2018a · 2018
Closest in time.
Cross-project code clones in GitHub
Mohammad Gharehyazie, Baishakhi Ray, Mehdi Keshani, Masoumeh Soleimani Zavosht, Abbas Heydarnoori, and Vladimir Filkov. 2018 · 2018
Closest in time.
A Retrieve-and-Edit Framework for Predicting Structured Outputs. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS)
Tatsunori B Hashimoto, Kelvin Guu, Yonatan Oren, and Percy Liang. 2018 · 2018
Closest in time.
Deep learning type inference. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . ACM, 152–162
Vincent J Hellendoorn, Christian Bird, Earl T Barr, and Miltiadis Allamanis. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PHOG: Probabilistic Model for Code. In Proceedings of the International Conference on Machine Learning (ICML) . 2933–2942
Pavol Bielik, Veselin Raychev, and Martin Vechev. 2016 · 2016
Cited alongside, same era.
Learning programs from noisy data. In Proceedings of the Symposium on Principles of Programming Languages (POPL) , Vol. 51. ACM, 761–774
Veselin Raychev, Pavol Bielik, Martin Vechev, and Andreas Krause. 2016 · 2016
Cited alongside, same era.
SourcererCC: scaling code clone detection to big-code. In Software Engineering (ICSE), 2016 IEEE/ACM 38th International Conference on . IEEE, 1157–1168
Hitesh Sajnani, Vaibhav Saini, Jeffrey Svajlenko, Chanchal K Roy, and Cristina V Lopes. 2016 · 2016
Cited alongside, same era.
A Parallel Corpus of Python Functions and Documentation Strings for Automated Code Documentation and Code Generation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers) , Vol. 2. 314–319
Antonio Valerio Miceli Barone and Rico Sennrich. 2017 · 2017
Cited alongside, same era.
Are deep neural networks the best choice for modeling source code?. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering . ACM, 763–773
Vincent J Hellendoorn and Premkumar Devanbu. 2017 · 2017
Cited alongside, same era.
Learning to Represent Programs with Graphs. In Proceedings of the International Conference on Learning Representations (ICLR)
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi. 2018b
Cited in the paper.
Closest in time.
Mapping Language to Code in Programmatic Context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . 1643–1652
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018 · 2018
Closest in time.
code2vec: Learning distributed representations of code
Uri Alon, Meital Zilberstein, Omer Levy, and Eran Yahav. 2019 · 2019
Closest in time.
Summarizing source code using a neural attention model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Vol. 1. 2073–2083
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016 · 2083
Closest in time.
A convolutional attention network for extreme summarization of source code. In Proceedings of the International Conference on Machine Learning (ICML) . 2091–2100
Miltiadis Allamanis, Hao Peng, and Charles Sutton. 2016 · 2091
Closest in time.