Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have become integral to various software engineering tasks, including code generation, bug detection, and repair.
A. Broder, “On the resemblance and containment of documents,” in Compression and Complexity of SEQUENCES 1997 (Cat. No.97TB100171) , 1997, pp. 21–29
1997
Earlier work this paper cites.
A. Gionis, P. Indyk, and R. Motwani, “Similarity search in high dimensions via hashing,” in International Conference on Very Large Data Bases , ser. VLDB ’99, 1999, p. 518–529
1999
Earlier work this paper cites.
M. Gabel and Z. Su, “A study of the uniqueness of source code,” in International Symposium on Foundations of Software Engineering , G. Roman and A. van der Hoek, Eds. ACM, 2010, pp. 147–156
2010
Earlier work this paper cites.
R. Just, D. Jalali, and M. D. Ernst, “Defects4j: a database of existing faults to enable controlled testing studies for java programs,” in International Symposium on Software Testing and Analysis , 2014, pp. 437–440
2014
Earlier work this paper cites.
W. E. Wong et al. , “A survey on software fault localization,” IEEE Transactions on Software Engineering , vol. 42, no. 8, pp. 707–740, 2016
2016
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “Squad: 100, 000+ questions for machine comprehension of text,” in Conference on Empirical Methods in Natural Language Processing , 2016, pp. 2383–2392. [Online]. Available: https://doi.org/10.18653/v1/d16-1264
2016
Earlier work this paper cites.
C. Le Goues, M. Pradel, and A. Roychoudhury, “Automated Program Repair,” Communications of the ACM , vol. 62, no. 12, pp. 56–65, 2019
2019
Earlier work this paper cites.
C. Zuo, Z. Lin, and Y. Zhang, “Why does your data leak? uncovering the data leakage in cloud from mobile apps,” in Symposium on Security and Privacy , 2019, pp. 1296–1310
2019
Earlier work this paper cites.
R. Widyasari et al. , “Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,” in Foundations of Software Engineering , 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks et al. , “Measuring massive multitask language understanding,” in International Conference on Learning Representations1 . OpenReview.net, 2021. [Online]. Available: https://openreview.net/forum?id=d7KBjmI3GmQ
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks et al. , “Measuring coding challenge competence with APPS,” in Neural Information Processing Systems Track on Datasets and Benchmarks , J. Vanschoren and S. Yeung, Eds., 2021
2021
Earlier work this paper cites.
2022
Cited alongside, same era.
K. Tirumala, A. H. Markosyan, L. Zettlemoyer, and A. Aghajanyan, “Memorization without overfitting: Analyzing the training dynamics of large language models,” in Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
G. An et al. , “Bugsc++: A highly usable real world defect benchmark for c/c++,” in International Conference on Automated Software Engineering . IEEE Press, 2024, p. 2034–2037. [Online]. Available: https://doi.org/10.1109/ASE56229.2023.00208
2024
Closest in time.
2024
Closest in time.
A. Silva, N. Saavedra, and M. Monperrus, “Gitbug-java: A reproducible benchmark of recent java bugs,” in International Conference on Mining Software Repositories , ser. MSR ’24. ACM, 2024, p. 118–122
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
E. Nijkamp et al. , “Codegen: An open large language model for code with multi-turn program synthesis,” in International Conference on Learning Representations . OpenReview.net, 2023
2023
Cited alongside, same era.
B. Rozière et al. , “Code llama: Open foundation models for code,” 2023. [Online]. Available: https://doi.org/10.48550/arXiv.2308.12950
2023
Cited alongside, same era.
H. Touvron et al. , “Llama: Open and efficient foundation language models,” 2023. [Online]. Available: https://doi.org/10.48550/arXiv.2302.13971
2023
Cited alongside, same era.
A. Q. Jiang et al. , “Mistral 7b,” 2023. [Online]. Available: https://arxiv.org/abs/2310.06825
2023
Cited alongside, same era.
A. Silva, S. Fang, and M. Monperrus, “Repairllama: Efficient representations and fine-tuned adapters for program repair,” CoRR , 2023. [Online]. Available: https://doi.org/10.48550/arXiv.2312.15698
2023
Cited alongside, same era.
D. Kocetkov et al. , “The stack: 3 TB of permissively licensed source code,” Trans. Mach. Learn. Res. , 2023. [Online]. Available: https://openreview.net/forum?id=pxpbTdUEpD
2023
Cited alongside, same era.
2023
Cited alongside, same era.
C. S. Xia and L. Zhang, “Keep the conversation going: Fixing 162 out of 337 bugs for $0.42 each using chatgpt,” CoRR , 2023. [Online]. Available: https://doi.org/10.48550/arXiv.2304.00385
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
A. Lozhkov et al. , “Starcoder 2 and the stack v2: The next generation,” 2024
2024
Closest in time.
G. Team et al. , “Gemma 2: Improving open language models at a practical size,” 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2408.00118
2024
Closest in time.
C. Team et al. , “Codegemma: Open code models based on gemma,” 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2406.11409
2024
Closest in time.
A. Z. H. Yang, C. Le Goues, R. Martins, and V. J. Hellendoorn, “Large language models for test-free fault localization,” in International Conference on Software Engineering . ACM, 2024, pp. 17:1–17:12
2024
Closest in time.
X. Zhou, T. Zhang, and D. Lo, “Large language model for vulnerability detection: Emerging results and future directions,” in International Conference on Software Engineering: New Ideas and Emerging Results . ACM, 2024, pp. 47–51
2024
Closest in time.
W. Shi et al. , “Detecting pretraining data from large language models,” in International Conference on Learning Representations . OpenReview.net, 2024. [Online]. Available: https://openreview.net/forum?id=zWqr3MQuNs
2024
Closest in time.