Fetching the paper…
Reading the bibliography…
Learning-based techniques, especially advanced Large Language Models (LLMs) for code, have gained considerable popularity in various software engineering (SE) tasks.
C. Gini, “Variabilit‡ e mutabilit,” Reprinted in Memorie di metodologica statistica (Ed. Pizetti E , 1912
1912
Earlier work this paper cites.
S. Hanson and L. Pratt, “Comparing biases for minimal network construction with back-propagation,” Advances in neural information processing systems , vol. 1, 1988
1988
Earlier work this paper cites.
C. Zhu, R. H. Byrd, P. Lu, and J. Nocedal, “Algorithm 778: L-bfgs-b: Fortran subroutines for large-scale bound-constrained optimization,” ACM Transactions on mathematical software (TOMS) , vol. 23, no. 4, pp. 550–560, 1997
1997
Earlier work this paper cites.
N. E. Fenton and N. Ohlsson, “Quantitative analysis of faults and failures in a complex software system,” IEEE Trans. Software Eng. , vol. 26, no. 8, pp. 797–814, 2000
2000
Earlier work this paper cites.
K. Crammer and Y. Singer, “On the algorithmic implementation of multiclass kernel-based vector machines,” Journal of machine learning research , vol. 2, no. Dec, pp. 265–292, 2001
2001
Earlier work this paper cites.
L. Breiman, “Random forests,” Machine learning , vol. 45, pp. 5–32, 2001
2001
Earlier work this paper cites.
C. Shirky, “Power laws, weblogs, and inequality,” 2003
2003
Earlier work this paper cites.
E. Brynjolfsson, Y. Hu, and M. D. Smith, “Consumer surplus in the digital economy: Estimating the value of increased product variety at online booksellers,” Management science , vol. 49, no. 11, pp. 1580–1596, 2003
2003
Earlier work this paper cites.
M. Dowd, J. McDonald, and J. Schuh, The art of software security assessment: Identifying and preventing software vulnerabilities . Pearson Education, 2006
2006
Earlier work this paper cites.
C. Anderson, The long tail: Why the future of business is selling less of more . Hachette UK, 2006
2006
Earlier work this paper cites.
D. Turner, M. Fossi, E. Johnson, T. Mack, J. Blackbird, S. Entwisle, M. K. Low, D. McKinney, and C. Wueest, “Symantec global internet security threat report trends for July–December 07, volume xii, april,” pp. 1–36, 2008
2008
Earlier work this paper cites.
T. Hastie, R. Tibshirani, J. H. Friedman, and J. H. Friedman, The elements of statistical learning: data mining, inference, and prediction . Springer, 2009, vol. 2
2009
Earlier work this paper cites.
G. Elahi, E. Yu, and N. Zannone, “A vulnerability-centric requirements engineering framework: analyzing security attacks, countermeasures, and requirements based on vulnerabilities,” Requirements engineering , vol. 15, pp. 41–62, 2010
2010
Earlier work this paper cites.
A. Hindle, E. T. Barr, Z. Su, M. Gabel, and P. T. Devanbu, “On the naturalness of software,” in 34th International Conference on Software Engineering, ICSE 2012, June 2-9, 2012, Zurich, Switzerland . IEEE Computer Society, 2012, pp. 837–847
2012
Earlier work this paper cites.
Y. Ko, “A study of term weighting schemes using class information for text classification,” in Proceedings of the 35th international ACM SIGIR conference on Research and development in information retrieval , 2012, pp. 1029–1030
2012
Earlier work this paper cites.
S. Wang and X. Yao, “Using class imbalance learning for software defect prediction,” IEEE Trans. Reliab. , vol. 62, no. 2, pp. 434–443, 2013
2013
Earlier work this paper cites.
B. Johnson, Y. Song, E. Murphy-Hill, and R. Bowdidge, “Why don’t software developers use static analysis tools to find bugs?” in 2013 35th International Conference on Software Engineering (ICSE) . IEEE, 2013, pp. 672–681
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
C. V. Lopes and J. Ossher, “How scale affects structure in java programs,” in Proceedings of the 2015 ACM SIGPLAN International Conference on Object-Oriented Programming, Systems, Languages, and Applications, OOPSLA 2015, part of SPLASH 2015, Pittsburgh, PA, USA, October 25-30, 2015 . ACM, 2015, pp. 675–694
2015
Earlier work this paper cites.
M. Tan, L. Tan, S. Dara, and C. Mayeux, “Online defect prediction for imbalanced data,” in 37th IEEE/ACM International Conference on Software Engineering, ICSE 2015, Florence, Italy, May 16-24, 2015, Volume 2 . IEEE Computer Society, 2015, pp. 99–108
2015
Earlier work this paper cites.
H. Borges, A. Hora, and M. T. Valente, “Understanding the factors that impact the popularity of github repositories,” in 2016 IEEE international conference on software maintenance and evolution (ICSME) . IEEE, 2016, pp. 334–344
2016
Earlier work this paper cites.
Y. Kamei, T. Fukushima, S. McIntosh, K. Yamashita, N. Ubayashi, and A. E. Hassan, “Studying just-in-time defect prediction using cross-project models,” Empir. Softw. Eng. , vol. 21, no. 5, pp. 2072–2106, 2016
2016
Earlier work this paper cites.
X. Gu, H. Zhang, D. Zhang, and S. Kim, “Deep api learning,” in Proceedings of the 2016 24th ACM SIGSOFT international symposium on foundations of software engineering , 2016, pp. 631–642
2016
Earlier work this paper cites.
J. Fowkes and C. Sutton, “Parameter-free probabilistic api mining across github,” in Proceedings of the 2016 24th ACM SIGSOFT international symposium on foundations of software engineering , 2016, pp. 254–265
2016
Earlier work this paper cites.
P. S. Kochhar, X. Xia, D. Lo, and S. Li, “Practitioners’ expectations on automated fault localization,” in Proceedings of the 25th International Symposium on Software Testing and Analysis , 2016, pp. 165–176
2016
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2980–2988
2017
Cited alongside, same era.
M. M. Öztürk, “Which type of metrics are useful to deal with class imbalance in software defect prediction?” Inf. Softw. Technol. , vol. 92, pp. 17–29, 2017
2017
Cited alongside, same era.
C. Tan, F. Sun, T. Kong, W. Zhang, C. Yang, and C. Liu, “A survey on deep transfer learning,” in Artificial Neural Networks and Machine Learning–ICANN 2018: 27th International Conference on Artificial Neural Networks, Rhodes, Greece, October 4-7, 2018, Proceedings, Part III 27 . Springer, 2018, pp. 270–279
2018
Cited alongside, same era.
K. E. Bennin, J. Keung, P. Phannachitta, A. Monden, and S. Mensah, “MAHAKIL: diversity based oversampling approach to alleviate the class imbalance issue in software defect prediction,” IEEE Trans. Software Eng. , vol. 44, no. 6, pp. 534–550, 2018
2018
X. Zhou, D. Han, and D. Lo, “Assessing generalizability of codebert,” in 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 2021, pp. 425–436
2021
Later among the works it cites.
S. Feng, J. Keung, X. Yu, Y. Xiao, K. E. Bennin, M. A. Kabir, and M. Zhang, “COSTE: complexity-based oversampling technique to alleviate the class imbalance problem in software defect prediction,” Inf. Softw. Technol. , vol. 129, p. 106432, 2021
2021
Later among the works it cites.
O. Dabic, E. Aghajani, and G. Bavota, “Sampling projects in github for msr studies,” in 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR) . IEEE, 2021, pp. 560–564
2021
Later among the works it cites.
J. Zhou, M. Pacheco, Z. Wan, X. Xia, D. Lo, Y. Wang, and A. E. Hassan, “Finding a needle in a haystack: Automated mining of silent vulnerability fixes,” in 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 2021, pp. 705–716
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie, “The inaturalist species classification and detection dataset,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 8769–8778
2018
Cited alongside, same era.
P. Martins, R. Achar, and C. V. Lopes, “50k-c: A dataset of compilable, and compiled, java projects,” in Proceedings of the 15th international conference on mining software repositories , 2018, pp. 1–5
2018
Cited alongside, same era.
F. Ebert, F. Castor, N. Novielli, and A. Serebrenik, “Communicative intention in code review questions,” in 2018 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 2018, pp. 519–523
2018
Cited alongside, same era.
X.-B. D. Le, F. Thung, D. Lo, and C. L. Goues, “Overfitting in semantics-based automated program repair,” in Proceedings of the 40th International Conference on Software Engineering , 2018, pp. 163–163
2018
Cited alongside, same era.
Z. Liu, Z. Miao, X. Zhan, J. Wang, B. Gong, and S. X. Yu, “Large-scale long-tailed recognition in an open world,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 2537–2546
2019
Cited alongside, same era.
G. G. Cabral, L. L. Minku, E. Shihab, and S. Mujahid, “Class imbalance evolution and verification latency in just-in-time software defect prediction,” in Proceedings of the 41st International Conference on Software Engineering, ICSE 2019, Montreal, QC, Canada, May 25-31, 2019 . IEEE / ACM, 2019, pp. 666–676
2019
Cited alongside, same era.
T. Hoang, H. K. Dam, Y. Kamei, D. Lo, and N. Ubayashi, “Deepjit: an end-to-end deep learning framework for just-in-time defect prediction,” in Proceedings of the 16th International Conference on Mining Software Repositories, MSR 2019, 26-27 May 2019, Montreal, Canada . IEEE / ACM, 2019, pp. 34–45
2019
Cited alongside, same era.
Y. Zhou, S. Liu, J. Siow, X. Du, and Y. Liu, “Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
Later among the works it cites.
Y. Zhou, J. K. Siow, C. Wang, S. Liu, and Y. Liu, “Spi: Automated identification of security patches via commits,” ACM Transactions on Software Engineering and Methodology (TOSEM) , vol. 31, no. 1, pp. 1–27, 2021
2021
Later among the works it cites.
R. Tufano, L. Pascarella, M. Tufano, D. Poshyvanyk, and G. Bavota, “Towards automating code review activities,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 2021, pp. 163–174
2021
Later among the works it cites.
2021
Later among the works it cites.
R. Tufano, S. Masiero, A. Mastropaolo, L. Pascarella, D. Poshyvanyk, and G. Bavota, “Using pre-trained models to boost code review automation,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 2291–2302
2022
Later among the works it cites.
L. Yang, H. Jiang, Q. Song, and J. Guo, “A survey on long-tailed visual recognition,” International Journal of Computer Vision , vol. 130, no. 7, pp. 1837–1872, 2022
2022
Later among the works it cites.
“2022 CWE Top 25 Most Dangerous Software Weaknesses.” https://cwe.mitre.org/top25/archive/2022/2022_cwe_top25.html , 2022
2022
Later among the works it cites.
C. Watson, N. Cooper, D. N. Palacio, K. Moran, and D. Poshyvanyk, “A systematic literature review on the use of deep learning in software engineering research,” ACM Transactions on Software Engineering and Methodology (TOSEM) , vol. 31, no. 2, pp. 1–58, 2022
2022
Later among the works it cites.
Z. Yang, J. Shi, J. He, and D. Lo, “Natural attack for pre-trained models of code,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 1482–1493
2022
Later among the works it cites.
J. A. H. López, M. Weyssow, J. S. Cuadrado, and H. A. Sahraoui, “Ast-probe: Recovering abstract syntax trees from hidden representations of pre-trained language models,” in 37th IEEE/ACM International Conference on Automated Software Engineering, ASE 2022, Rochester, MI, USA, October 10-14, 2022 . ACM, 2022, pp. 11:1–11:11
2022
Later among the works it cites.
J. Shi, Z. Yang, B. Xu, H. J. Kang, and D. Lo, “Compressing pre-trained models of code into 3 MB,” in 37th IEEE/ACM International Conference on Automated Software Engineering, ASE 2022, Rochester, MI, USA, October 10-14, 2022 . ACM, 2022, pp. 24:1–24:12
2022
Later among the works it cites.
S. Alshammari, Y.-X. Wang, D. Ramanan, and S. Kong, “Long-tailed recognition via weight balancing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 6897–6907
2022
Later among the works it cites.
J. Martin and J. L. Guo, “Deep api learning revisited,” in Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension , 2022, pp. 321–330
2022
Later among the works it cites.
P. Thongtanunam, C. Pornprasit, and C. Tantithamthavorn, “Autotransform: Automated code transformation to support modern code review process,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) , 2022
2022
Later among the works it cites.
V.-A. Nguyen, D. Q. Nguyen, V. Nguyen, T. Le, Q. H. Tran, and D. Phung, “Regvd: Revisiting graph neural networks for vulnerability detection,” in Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings , 2022, pp. 178–182
2022
Later among the works it cites.
E. Shi, Y. Wang, W. Tao, L. Du, H. Zhang, S. Han, D. Zhang, and H. Sun, “Race: Retrieval-augmented commit message generation,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 5520–5530
2022
Later among the works it cites.
I. C. Irsan, T. Zhang, F. Thung, K. Kim, and D. Lo, “Multi-modal api recommendation,” in 30th IEEE International Conference on Software Analysis, Evolution and Reengineering, SANER 2023, Macao SAR, China, March 21st-24th, 2023 , 2023
2023
Closest in time.
S. Pan, L. Bao, X. Xia, D. Lo, and S. Li, “Fine-grained commit-level vulnerability type prediction by CWE tree structure,” in 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, 2023, pp. 957–969
2023
Closest in time.
W. Jiang, N. Synovic, M. Hyatt, T. R. Schorlemmer, R. Sethi, Y. Lu, G. K. Thiruvathukal, and J. C. Davis, “An empirical study of pre-trained model reuse in the hugging face deep learning model registry,” in 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, 2023, pp. 2463–2475
2023
Closest in time.
C. Niu, C. Li, V. Ng, D. Chen, J. Ge, and B. Luo, “An empirical comparison of pre-trained models of source code,” in 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, 2023, pp. 2136–2148
2023
Closest in time.
S. Liu, B. Wu, X. Xie, G. Meng, and Y. Liu, “Contrabert: Enhancing code pre-trained models via contrastive learning,” in 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, 2023, pp. 2476–2487
2023
Closest in time.
R. Croft, M. A. Babar, and M. M. Kholoosi, “Data quality for software vulnerability datasets,” in 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, 2023, pp. 121–133
2023
Closest in time.
“difflib, a library to extract the token-level edits.” https://docs.python.org/3.7/library/difflib.html , 2023
2023
Closest in time.