Fetching the paper…
Reading the bibliography…
Data augmentation has become a standard practice in software engineering to address limited or imbalanced data sets, particularly in specialized domains like test classification and bug detection where data can be scarce.
N. Chawla, K. Bowyer, L. Hall, and W. Kegelmeyer, “SMOTE: Synthetic minority over-sampling technique,” J. Artif. Intell. Res. (JAIR) , vol. 16, pp. 321–357, 06 2002
2002
Earlier work this paper cites.
B. Jahić, N. Guelfi, and B. Ries, “Software engineering for dataset augmentation using generative adversarial networks,” in 2019 IEEE 10th International Conference on Software Engineering and Service Science (ICSESS) , 2019, pp. 59–66
2019
Earlier work this paper cites.
W. Lam, P. Godefroid, S. Nath, A. Santhiar, and S. Thummalapenta, “Root causing flaky tests in a large-scale industrial setting,” in Proc. of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2019) , 2019, p. 101–111
2019
Earlier work this paper cites.
Z. Feng, D. Guo, D. Tang et al. , “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , Nov. 2020, pp. 1536–1547
2020
Earlier work this paper cites.
Y. Yu, Y. Zhuang, J. Zhang, Y. Meng, A. J. Ratner, R. Krishna, J. Shen, and C. Zhang, “Large language model as attributed training data generator: A tale of diversity and bias,” Advances in Neural Information Processing Systems , vol. 36, pp. 55 734–55 784, 2023
2023
Cited alongside, same era.
A. Akli, G. Haben, S. Habchi, M. Papadakis, and Y. Le Traon, “FlakyCat: Predicting flaky tests categories using few-shot learning,” in Proc. of IEEE/ACM International Conference on Automation of Software Test (AST 2023) , May 2023, pp. 140–151
2023
Cited alongside, same era.
Y. Zhou, C. Guo, X. Wang, Y. Chang, and Y. Wu, “A survey on data augmentation in large model era,” arXiv preprint:2401.15422 , 2024
2024
Cited alongside, same era.
B. Ding, C. Qin, R. Zhao, T. Luo, X. Li, G. Chen, W. Xia, J. Hu, A. T. Luu, and S. Joty, “Data augmentation using LLMs: Data perspectives, learning paradigms and challenges,” in Findings of the Association for Computational Linguistics: ACL 2024 , Aug. 2024, pp. 1679–1705
Y. He, J. Wang, Y. Rong, and H. Chen, “FuzzAug: Data augmentation by fuzzing for neural test generation,” arXiv preprint:2406.08665 , 2024
2024
Later among the works it cites.
T. H. M. Le and M. Ali Babar, “Mitigating data imbalance for software vulnerability assessment: Does data augmentation help?” in Proc. of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM 2024) , 2024, pp. 119–130
2024
Later among the works it cites.
R. More and J. S. Bradbury, “An analysis of LLM fine-tuning and few-shot learning for flaky test detection and classification,” in Proc. of the 18th IEEE International Conference on Software Testing, Verification and Validation (ICST 2025) , Apr. 2025
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.