Fetching the paper…
Reading the bibliography…
This paper introduces a novel code-to-code search technique that enhances the performance of Large Language Models (LLMs) by including both static and dynamic features as well as utilizing both similar and dissimilar examples during training.
K. Zhang, “A simple algorithm for tree edit distance,” in Proceedings of the 1st annual ACM-SIAM symposium on Discrete algorithms . Society for Industrial and Applied Mathematics, 1989, pp. 90–97
1989
Earlier work this paper cites.
P. Willett, “Some issues in the evaluation of text retrieval systems,” in Proceedings of the seventeenth annual international ACM SIGIR conference on Research and development in information retrieval . ACM, 1994, pp. 188–195
1994
Earlier work this paper cites.
I. D. Baxter, A. Yahin, L. Moura, M. Sant’Anna, and L. Bier, “Clone detection using abstract syntax trees,” in Proceedings. International Conference on Software Maintenance (Cat. No. 98CB36272) . IEEE, 1998, pp. 368–377
1998
Earlier work this paper cites.
T. Kamiya, S. Kusumoto, and K. Inoue, “Ccfinder: a multilinguistic token-based code clone detection system for large scale source code,” IEEE Transactions on Software Engineering , vol. 28, no. 7, pp. 654–670, 2002
2002
Earlier work this paper cites.
C. M. Bishop, “Pattern recognition and machine learning,” Springer , 2006
2006
Earlier work this paper cites.
Z. Li, S. Lu, S. Myagmar, and Y. Zhou, “Cp-miner: finding copy-paste and related bugs in large-scale software code,” IEEE Transactions on Software Engineering , vol. 32, no. 3, pp. 176–192, 2006
2006
Earlier work this paper cites.
A. Moffat and J. Zobel, “Average rank performance metrics,” Journal of Information Retrieval , vol. 11, no. 4, pp. 331–353, 2008
2008
Earlier work this paper cites.
C. S. Marcum, J. J. Haas, and J. Clause, “A comparative study of static and dynamic code analysis for bug detection,” in 2010 IEEE International Conference on Software Maintenance . IEEE, 2010, pp. 1–10
2010
Earlier work this paper cites.
D. Kim, J. Nam, and J.-G. Song, “Automatic patch generation learned from human-written patches,” in Proceedings of the 35th International Conference on Software Engineering . IEEE/ACM, 2013, pp. 802–811
2013
Earlier work this paper cites.
Y. Kamei, A. Monden, and K.-i. Matsumoto, “Cross-language program search using a multilingual abstract syntax tree,” in Proceedings of the 2013 IEEE International Conference on Software Maintenance . IEEE, 2013, pp. 320–329
2013
Earlier work this paper cites.
A. Rastogi, R. S. Gu, K. Kumar, and S. Chandra, “Prose: Programming language semantics for large-scale code search,” in Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data . ACM, 2015, pp. 1745–1757
2015
Earlier work this paper cites.
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller, “Action-conditional video prediction using deep networks in atari games,” in Advances in Neural Information Processing Systems , 2015, pp. 2863–2871
2015
Earlier work this paper cites.
H. Pham, T. Tran, and S. Venkatesh, “Modelling dynamic textual content with temporal language models,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 6, no. 3, p. 33, 2015
2015
Earlier work this paper cites.
B. Ray, D. Kim, B. Vasilescu, S. Godhane, and V. Filkov, “Onward! toward more realistic program benchmarks.” ACM, 2016, pp. 88–103
2016
Earlier work this paper cites.
F.-H. Su, J. Bell, G. Kaiser, and S. Sethumadhavan, “Identifying Functionally Similar Code in Complex Codebases,” in 2016 IEEE 24th International Conference on Program Comprehension (ICPC) , 2016, pp. 1–10
2016
Earlier work this paper cites.
L. Wu, T. Chen, and Y. Xie, “Cross-language code search with a bilingual semantic space,” in Proceedings of the 38th International Conference on Software Engineering Companion . ACM, 2016, pp. 13–22
2016
Earlier work this paper cites.
F.-H. Su, J. Bell, K. Harvey, S. Sethumadhavan, G. Kaiser, and T. Jebara, “Code Relatives: Detecting Similarly Behaving Software,” ser. FSE 2016. New York, NY, USA: Association for Computing Machinery, 2016, p. 702–714. [Online]. Available: https://doi.org/10.1145/2950290.2950321
2016
Earlier work this paper cites.
H. Sajnani, V. Saini, J. Svajlenko, C. K. Roy, and C. V. Lopes, “Sourcerercc: Scaling code clone detection to big-code,” ser. ICSE ’16. New York, NY, USA: Association for Computing Machinery, 2016, p. 1157–1168. [Online]. Available: https://doi.org/10.1145/2884781.2884877
2016
Earlier work this paper cites.
B. He and J. Liu, “Deeplog: Anomaly detection and diagnosis from system logs through deep learning,” in 2016 46th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) . IEEE, 2016, pp. 230–241
2016
Cited alongside, same era.
D. Kim, J. Nam, and J.-G. Song, “Mining libraries and frameworks for api migration,” in 2017 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 2017, pp. 358–367
2017
Cited alongside, same era.
M. Zhou, Z. Zhang, J. Wu, H. Zhang, and H. Mei, “Automatic cross-language api adaptation via enhanced search-based api code transplantation,” in Proceedings of the 40th International Conference on Software Engineering: Companion Proceeedings , 2018, pp. 81–82
2018
Cited alongside, same era.
F. Long and M. Rinard, “Cross-language program repair using conditioned symbolic execution,” in 2018 IEEE/ACM 40th International Conference on Software Engineering (ICSE) . IEEE, 2018, pp. 738–748
2018
G. Mathew, C. Parnin, and K. T. Stolee, “Slacc: Simion-based language agnostic code clones,” in 2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE) , 2020, pp. 210–221
2020
Later among the works it cites.
G. Mathew, C. Parnin, and K. T. Stolee, “Slacc: Simion-based language agnostic code clones,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering , 2020, pp. 210–221
2020
Later among the works it cites.
Microsoft, “Github,” (Accessed on May 01 2020). [Online]. Available: www.github.com
2020
Later among the works it cites.
T. Yang and R. Fonseca, “Practical runtime specialization for higher-order functions,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21) , 2021, pp. 601–614
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
R. Yue, Z. Gao, N. Meng, Y. Xiong, X. Wang, and J. D. Morgenthaler, “Automatic clone recommendation for refactoring based on the present and the past,” in 2018 IEEE International Conference on Software Maintenance and Evolution (ICSME) , 2018, pp. 115–126
2018
Cited alongside, same era.
R. Xia, X. Xia, and J. Sun, “A neural model for cross-language code retrieval,” in Proceedings of the 40th International Conference on Software Engineering . ACM, 2018, pp. 245–256
2018
Cited alongside, same era.
K. Kim, D. Kim, T. F. Bissyandé, E. Choi, L. Li, J. Klein, and Y. L. Traon, “Facoy: A code-to-code search engine,” ser. ICSE ’18. New York, NY, USA: Association for Computing Machinery, 2018, p. 946–957. [Online]. Available: https://doi.org/10.1145/3180155.3180187
2018
Cited alongside, same era.
F.-H. Su, J. Bell, G. Kaiser, and B. Ray, “Obfuscation Resilient Search through Executable Classification,” ser. MAPL 2018. New York, NY, USA: Association for Computing Machinery, 2018, p. 20–30. [Online]. Available: https://doi.org/10.1145/3211346.3211352
2018
Cited alongside, same era.
M. Allamanis, M. Brockschmidt, and M. Khademi, “Learning to represent programs with graphs,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
P. Feng, J. Ma, C. Sun, X. Xu, and Y. Ma, “A novel dynamic android malware detection system with ensemble learning,” IEEE Access , vol. 6, pp. 30 996–31 011, 2018
2018
Cited alongside, same era.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , J. Burstein, C. Doran, and T. Solorio, Eds. Association for Computational Linguistics, 2019, pp. 4171–4186. [Online]. Available: https://doi.org/10.18653/v1/n19-1423
2019
Cited alongside, same era.
B. Vasilescu, A. Serebrenik, and V. Filkov, “The state of the art in distributed software development: A systematic review,” in 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE) . IEEE, 2019, pp. 69–79
2019
Cited alongside, same era.
2021
Later among the works it cites.
K. Nakamura, R. Iwasaki, S. Hoshino, and I. Sato, “Atcoder: A dataset for machine learning on programming,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . ACM, 2021, pp. 3601–3611
2021
Later among the works it cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” 2021
2021
Later among the works it cites.
V. Nitin, A. Saieva, B. Ray, and G. Kaiser, “Direct: A transformer-based model for decompiled variable name recovery,” NLP4Prog 2021 , p. 48, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Zou, B. Ban, Y. Xue, and Y. Xu, “Ccgraph: A pdg-based code clone detector with approximate graph matching,” ser. ASE ’20. New York, NY, USA: Association for Computing Machinery, 2021, p. 931–942. [Online]. Available: https://doi.org/10.1145/3324884.3416541
2021
Later among the works it cites.
D. Guo, S. Ren, S. Lu, Z. Feng, D. Tang, S. Liu, L. Zhou, N. Duan, J. Yin, D. Jiang et al. , “Graphcodebert: Pre-training code representations with data flow,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
2022
Later among the works it cites.
K. Pei, Z. Xuan, J. Yang, S. Jana, and B. Ray, “Learning Approximate Execution Semantics From Traces for Binary Function Similarity,” IEEE Transactions on Software Engineering , vol. 49, no. 4, pp. 2776–2790, April 2023. [Online]. Available: https://doi.org/10.1109/TSE.2022.3231621
2022
Later among the works it cites.
Y. Ding, L. Buratti, S. Chakraborty, S. Pujar, A. Morari, and B. Ray, “Contrastive learning for source code with structural and functional properties,” 2022. [Online]. Available: https://openreview.net/forum?id=7KgeqhkbZab
2022
Later among the works it cites.
OpenAI, “Openai’s text embeddings,” 2023. [Online]. Available: https://platform.openai.com/docs/guides/embeddings
2023
Closest in time.
L. Di Grazia and M. Pradel, “Code search: A survey of techniques for finding code,” ACM Comput. Surv. , vol. 55, no. 11, feb 2023. [Online]. Available: https://doi.org/10.1145/3565971
2023
Closest in time.
S. N. Pinku, D. Mondal, and C. K. Roy, “Pathways to leverage transcompiler based data augmentation for cross-language clone detection,” 2023
2023
Closest in time.