Fetching the paper…
Reading the bibliography…
Much of software-engineering research relies on the naturalness of code, the fact that code, in small code snippets, is repetitive and can be predicted using statistical language models like n-gram.
——, “A mathematical theory of communication,” The Bell System Technical Journal , vol. 27, no. 3, pp. 379–423, 1948
1948
Earlier work this paper cites.
C. E. Shannon, “Prediction and entropy of printed english,” The Bell System Technical Journal , vol. 30, no. 1, pp. 50–64, 1951
1951
Earlier work this paper cites.
R. Kneser and H. Ney, “Improved backing-off for m-gram language modeling,” in 1995 International Conference on Acoustics, Speech, and Signal Processing , vol. 1, 1995, pp. 181–184 vol.1
1995
Earlier work this paper cites.
S. F. Chen and J. Goodman, “An empirical study of smoothing techniques for language modeling,” Computer Speech & Language , vol. 13, no. 4, pp. 359–394, 1999. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0885230899901286
1999
Earlier work this paper cites.
2002
Earlier work this paper cites.
T. Copeland, PMD applied . Centennial Books San Francisco, 2005, vol. 10
2005
Earlier work this paper cites.
D. Hovemeyer and W. Pugh, “Finding more null pointer bugs, but not too many,” in Proceedings of the 7th ACM SIGPLAN-SIGSOFT workshop on Program analysis for software tools and engineering , 2007, pp. 9–14
2007
Earlier work this paper cites.
A. Hindle, E. T. Barr, Z. Su, M. Gabel, and P. Devanbu, “On the naturalness of software,” in Proceedings of the 34th International Conference on Software Engineering , ser. ICSE ’12. IEEE Press, 2012, p. 837–847
2012
Earlier work this paper cites.
M. Allamanis and C. Sutton, “Mining source code repositories at massive scale using language modeling,” in 2013 10th working conference on mining software repositories (MSR) . IEEE, 2013, pp. 207–216
2013
Earlier work this paper cites.
V. J. Hellendoorn, P. T. Devanbu, and A. Bacchelli, “Will they like this? evaluating code contributions with language models,” in 2015 IEEE/ACM 12th Working Conference on Mining Software Repositories , 2015, pp. 157–167
2015
Cited alongside, same era.
R. Pawlak, M. Monperrus, N. Petitprez, C. Noguera, and L. Seinturier, “Spoon: A Library for Implementing Analyses and Transformations of Java Source Code,” Software: Practice and Experience , vol. 46, pp. 1155–1179, 2015. [Online]. Available: https://hal.archives-ouvertes.fr/hal-01078532/document
2015
Cited alongside, same era.
S. Wang, D. Chollak, D. Movshovitz-Attias, and L. Tan, “Bugram: Bug detection with n-gram language models,” in Proceedings of the 31st IEEE/ACM International Conference on Automated Software Engineering , ser. ASE 2016. New York, NY, USA: Association for Computing Machinery, 2016, p. 708–719. [Online]. Available: https://doi.org/10.1145/2970276.2970341
2016
Cited alongside, same era.
B. Lin, C. Nagy, G. Bavota, and M. Lanza, “On the impact of refactoring operations on code naturalness,” in 2019 IEEE 26th International Conference on Software Analysis, Evolution and Reengineering (SANER) , 2019, pp. 594–598
2019
Later among the works it cites.
M. Rahman, D. Palani, and P. C. Rigby, “Natural software revisited,” in 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE) , 2019, pp. 37–48
2019
Later among the works it cites.
T. T. Chekam, M. Papadakis, T. F. Bissyandé, Y. L. Traon, and K. Sen, “Selecting fault revealing mutants,” Empir. Softw. Eng. , vol. 25, no. 1, pp. 434–487, 2020. [Online]. Available: https://doi.org/10.1007/s10664-019-09778-7
2020
Later among the works it cites.
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “Codebert: A pre-trained model for programming and natural languages,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, EMNLP 2020, Online Event, 16-20 November 2020 , ser. Findings of ACL, T. Cohn, Y. He, and Y. Liu, Eds., vol. EMNLP 2020. Association for Computational Linguistics, 2020, pp. 1536–1547. [Online]. Available: https://doi.org/10.18653/v1/2020.findings-emnlp.139
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Ray, V. Hellendoorn, S. Godhane, Z. Tu, A. Bacchelli, and P. Devanbu, “On the ”naturalness” of buggy code,” in 2016 IEEE/ACM 38th International Conference on Software Engineering (ICSE) , 2016, pp. 428–439
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
M. Jimenez, C. Maxime, Y. Le Traon, and M. Papadakis, “On the impact of tokenizer and parameters on n-gram based code analysis,” in 2018 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 2018, pp. 437–448
2018
Cited alongside, same era.
2018
Cited alongside, same era.
M. Jimenez, C. Maxime, Y. Le Traon, and M. Papadakis, “Tuna: Tuning naturalness-based analysis,” in 2018 IEEE International Conference on Software Maintenance and Evolution (ICSME) , 2018, pp. 715–715
2018
Cited alongside, same era.
“Javaparser,” Available on https://github.com/javaparser/javaparser , https://javaparser.org/
Cited in the paper.
“Codebert,” https://github.com/microsoft/CodeBERT
Cited in the paper.
2020
Later among the works it cites.
M. K. Thota, F. H. Shajin, and P. Rajesh, “Survey on software defect prediction techniques,” International Journal of Applied Science and Engineering , vol. 17, pp. 331–344, December 2020
2020
Later among the works it cites.
T. Sharma, V. Efstathiou, P. Louridas, and D. Spinellis, “Code smell detection by deep direct-learning and transfer-learning,” Journal of Systems and Software , vol. 176, p. 110936, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0164121221000339
2021
Later among the works it cites.
D. Posnett, A. Hindle, and P. Devanbu, “Reflections on: A simpler model of software readability,” SIGSOFT Softw. Eng. Notes , vol. 46, no. 3, p. 30–32, jul 2021. [Online]. Available: https://doi.org/10.1145/3468744.3468754
2021
Later among the works it cites.
S. H. Alexander Trautsch, Fabian Trautsch, “The smartshark repository mining data,” 2021
2021
Later among the works it cites.
Z. Sun, J. M. Zhang, Y. Xiong, M. Harman, M. Papadakis, and L. Zhang, “Improving machine translation systems via isotopic replacement,” in 2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE) , 2022, pp. 1181–1192
2022
Closest in time.