Fetching the paper…
Reading the bibliography…
In this work, we revisit existing oracle generation studies plus ChatGPT to empirically investigate the current standing of their performance in both NLG-based and test adequacy metrics.
L. R. Dice, “Measures of the amount of ecologic association between species,” Ecology , vol. 26, no. 3, pp. 297–302, 1945
1945
Earlier work this paper cites.
T. T. Tanimoto, “Elementary mathematical theory of classification and prediction,” 1958
1958
Earlier work this paper cites.
S. S. Smith, “Scope and methods of political science,” Teaching Political Science , vol. 9, no. 1, pp. 60–62, 1981
1981
Earlier work this paper cites.
“Accuracy (trueness and precision) of measurement methods and results — part 1: General principles and definitions,” International Organization for Standardization, Geneva, CH, Standard, 1994
1994
Earlier work this paper cites.
D. Peters and D. L. Parnas, “Generating a test oracle from program documentation: work in progress,” in Proceedings of the 1994 ACM SIGSOFT international symposium on Software testing and analysis , 1994, pp. 58–65
1994
Earlier work this paper cites.
L. Du Bousquet, F. Ouabdesselam, J.-L. Richier, and N. Zuanon, “Lutess: a specification-driven testing environment for synchronous software,” in Proceedings of the 21st international conference on Software engineering , 1999, pp. 267–276
1999
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
M. M. Tikir and J. K. Hollingsworth, “Efficient instrumentation for code coverage testing,” ACM SIGSOFT Software Engineering Notes , vol. 27, no. 4, pp. 86–96, 2002
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out , 2004, pp. 74–81
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
C. Pacheco and M. D. Ernst, “Eclat: Automatic generation and classification of test inputs,” in ECOOP 2005-Object-Oriented Programming: 19th European Conference, Glasgow, UK, July 25-29, 2005. Proceedings 19 . Springer, 2005, pp. 504–527
2005
Earlier work this paper cites.
C. Pacheco and M. D. Ernst, “Randoop: feedback-directed random testing for java,” in Companion to the 22nd ACM SIGPLAN conference on Object-oriented programming systems and applications companion , 2007, pp. 815–816
2007
Earlier work this paper cites.
G. Fraser and A. Arcuri, “Evosuite: automatic test suite generation for object-oriented software,” in Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering , 2011, pp. 416–419
2011
Earlier work this paper cites.
S. R. Shahamiri, W. M. N. W. Kadir, S. Ibrahim, and S. Z. M. Hashim, “An automated framework for software test oracle,” Information and Software Technology , vol. 53, no. 7, pp. 774–788, 2011
2011
Earlier work this paper cites.
Y. Liu, Y. Zhou, S. Wen, and C. Tang, “A strategy on selecting performance metrics for classifier evaluation,” International Journal of Mobile Computing and Multimedia Communications (IJMCMC) , vol. 6, no. 4, pp. 20–35, 2014
2014
Cited alongside, same era.
Y. Graham, T. Baldwin, and N. Mathur, “Accurate evaluation of segment-level machine translation metrics,” in Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2015, pp. 1183–1191
2015
Cited alongside, same era.
H. Coles, T. Laurent, C. Henard, M. Papadakis, and A. Ventresque, “Pit: a practical mutation testing tool for java,” in Proceedings of the 25th international symposium on software testing and analysis , 2016, pp. 449–452
2016
Cited alongside, same era.
P. Liu, X. Zhang, M. Pistoia, Y. Zheng, M. Marques, and L. Zeng, “Automatic text input generation for mobile testing,” in 2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE) . IEEE, 2017, pp. 643–653
D. Roy, S. Fakhoury, and V. Arnaoudova, “Reassessing automatic evaluation metrics for code summarization tasks,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2021, pp. 1105–1116
2021
Later among the works it cites.
A. Shimorina, “Human vs automatic metrics: on the importance of correlation design,” 2021
2021
Later among the works it cites.
M. Tufano, S. K. Deng, N. Sundaresan, and A. Svyatkovskiy, “Methods2test: A dataset of focal methods mapped to test cases,” in Proceedings of the 19th International Conference on Mining Software Repositories . ACM, may 2022. [Online]. Available: https://doi.org/10.48550/arXiv.2203.12776
2022
Later among the works it cites.
H. Yu, Y. Lou, K. Sun, D. Ran, T. Xie, D. Hao, Y. Li, G. Li, and Q. Wang, “Automated assertion generation via information retrieval and its integration with deep learning.” ICSE, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
L. Saes, “Unit test generation using machine learning,” Universiteit van Amsterdamg , 2018
2018
Cited alongside, same era.
P. Schober, C. Boer, and L. A. Schwarte, “Correlation coefficients: appropriate use and interpretation,” Anesthesia & analgesia , vol. 126, no. 5, pp. 1763–1768, 2018
2018
Cited alongside, same era.
E. Reiter, “A structured review of the validity of bleu,” Computational Linguistics , vol. 44, no. 3, pp. 393–401, 2018
2018
Cited alongside, same era.
A. LeClair and C. McMillan, “Recommendations for datasets for source code summarization,” in Proceedings of NAACL-HLT , 2019, pp. 3931–3937
2019
Cited alongside, same era.
C. Watson, M. Tufano, K. Moran, G. Bavota, and D. Poshyvanyk, “On learning meaningful assert statements for unit test cases,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering , 2020, pp. 1398–1409
2020
Cited alongside, same era.
2020
Cited alongside, same era.
S. Stapleton, Y. Gambhir, A. LeClair, Z. Eberhart, W. Weimer, K. Leach, and Y. Huang, “A human study of comprehension and code summarization,” in Proceedings of the 28th International Conference on Program Comprehension , 2020, pp. 2–13
2020
Cited alongside, same era.
S. Liu and S. Nakajima, “Automatic test case and test oracle generation based on functional scenarios in formal specifications for conformance testing,” IEEE Transactions on Software Engineering , vol. 48, no. 2, pp. 691–712, 2020
2020
Cited alongside, same era.
E. Dinella, G. Ryan, T. Mytkowicz, and S. K. Lahiri, “Toga: A neural method for test oracle generation,” in Proceedings of the 44th International Conference on Software Engineering , ser. ICSE ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 2130–2141. [Online]. Available: https://doi.org/10.1145/3510003.3510141
2022
Later among the works it cites.
M. Tufano, D. Drain, A. Svyatkovskiy, and N. Sundaresan, “Generating accurate assert statements for unit test cases using pretrained transformers,” in Proceedings of the 3rd ACM/IEEE International Conference on Automation of Software Test , 2022, pp. 54–64
2022
Later among the works it cites.
Z. Yuan, Y. Lou, M. Liu, S. Ding, K. Wang, Y. Chen, and X. Peng, “No more manual tests? evaluating and improving chatgpt for unit test generation,” 2023
2023
Closest in time.
Wikipedia Contributors, “Overlap — Wikipedia, the free encyclopedia,” 2023, [Online; accessed 18-January-2023]. [Online]. Available: https://en.wikipedia.org/w/index.php?title=Overlap&oldid=1061948530
2023
Closest in time.
C. Lemieux, J. P. Inala, S. K. Lahiri, and S. Sen, “Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) , 2023
2023
Closest in time.
2023
Closest in time.
OpenAI. (2023) Chatgpt. [Online]. Available: https://openai.com/chatgpt
2023
Closest in time.
J. P. Lim and H. Lauw, “Large-scale correlation analysis of automated metrics for topic models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 13 874–13 898
2023
Closest in time.