Fetching the paper…
Reading the bibliography…
Large language models for code (i.e., code LLMs) have shown strong code understanding and generation capabilities.
H. Hemmati, “How effective are code coverage criteria?” in 2015 IEEE International Conference on Software Quality, Reliability and Security . IEEE, 2015, pp. 151–156
2015
Earlier work this paper cites.
P. Ammann and J. Offutt, Introduction to software testing . Cambridge University Press, 2016
2016
Earlier work this paper cites.
J. Henkel, S. K. Lahiri, B. Liblit, and T. Reps, “Code vectors: understanding programs through embedded abstracted symbolic traces,” in Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , ser. ESEC/FSE 2018. New York, NY, USA: Association for Computing Machinery, 2018, p. 163–174. [Online]. Available: https://doi.org/10.1145/3236024.3236085
2018
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang et al. , “Codexglue: A machine learning benchmark dataset for code understanding and generation,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1) , 2021
2021
Earlier work this paper cites.
Y. Elazar, N. Kassner, S. Ravfogel, A. Ravichander, E. Hovy, H. Schütze, and Y. Goldberg, “Measuring and improving consistency in pretrained language models,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1012–1031, 2021
2021
Earlier work this paper cites.
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba, “Evaluating large language models trained on code,” 2021
2021
Earlier work this paper cites.
F. Barbieri, L. E. Anke, and J. Camacho-Collados, “Xlm-t: Multilingual language models in twitter for sentiment analysis and beyond,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference , 2022, pp. 258–266
2022
Earlier work this paper cites.
A. Creswell, M. Shanahan, and I. Higgins, “Selection-inference: Exploiting large language models for interpretable logical reasoning,” in The Eleventh International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
F. Tsimpourlas, G. Rooijackers, A. Rajan, and M. Allamanis, “Embedding and classifying test execution traces using neural networks,” IET Software , vol. 16, no. 3, pp. 301–316, 2022. [Online]. Available: https://ietresearch.onlinelibrary.wiley.com/doi/abs/10.1049/sfw2.12038
2022
Earlier work this paper cites.
M. Jang, D. S. Kwon, and T. Lukasiewicz, “Becel: Benchmark for consistency evaluation of language models,” in Proceedings of the 29th International Conference on Computational Linguistics , 2022, pp. 3680–3696
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” in The Eleventh International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
B. Min, H. Ross, E. Sulem, A. P. B. Veyseh, T. H. Nguyen, O. Sainz, E. Agirre, I. Heintz, and D. Roth, “Recent advances in natural language processing via large pre-trained language models: A survey,” ACM Computing Surveys , vol. 56, no. 2, pp. 1–40, 2023
2023
Earlier work this paper cites.
A. Rogers, M. Gardner, and I. Augenstein, “Qa dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension,” ACM Computing Surveys , vol. 55, no. 10, pp. 1–45, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology , 2023
2023
Earlier work this paper cites.
A. Ni, S. Iyer, D. Radev, V. Stoyanov, W.-t. Yih, S. Wang, and X. V. Lin, “Lever: Learning to verify language-to-code generation with execution,” in International Conference on Machine Learning . PMLR, 2023, pp. 26 106–26 128
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. J. Min, Y. Ding, L. Buratti, S. Pujar, G. Kaiser, S. Jana, and B. Ray, “Beyond accuracy: Evaluating self-consistency of code llms,” in The Twelfth International Conference on Learning Representations , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
Y. Ding, B. Steenhoek, K. Pei, G. Kaiser, W. Le, and B. Ray, “Traced: Execution-aware pre-training for source code,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024, pp. 1–12
2024
Closest in time.
“GitHub Copilot · Your AI pair programmer,” https://github.com/features/copilot , 2024, last accessed Mar. 2024
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Jabbar, H. Hemmati, and R. Feldt, “Investigating execution trace embedding for test case prioritization,” in 2023 IEEE 23rd International Conference on Software Quality, Reliability, and Security (QRS) , 2023, pp. 279–290
2023
Cited alongside, same era.
M. Jang and T. Lukasiewicz, “Consistency analysis of chatgpt,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 15 970–15 985
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, W.-t. Yih, D. Fried, S. Wang, and T. Yu, “Ds-1000: A natural and reliable benchmark for data science code generation,” in International Conference on Machine Learning . PMLR, 2023, pp. 18 319–18 345
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
X. Du, M. Liu, K. Wang, H. Wang, J. Liu, Y. Chen, J. Feng, C. Sha, X. Peng, and Y. Lou, “Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Yu, B. Shen, D. Ran, J. Zhang, Q. Zhang, Y. Ma, G. Liang, Y. Li, Q. Wang, and T. Xie, “Codereval: A benchmark of pragmatic code generation with generative pre-trained models,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024, pp. 1–12
2024
Closest in time.
Y. Ding, Z. Wang, W. Ahmad, H. Ding, M. Tan, N. Jain, M. K. Ramanathan, R. Nallapati, P. Bhatia, D. Roth et al. , “Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
A. Z. Yang, C. Le Goues, R. Martins, and V. Hellendoorn, “Large language models for test-free fault localization,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024, pp. 1–12
2024
Closest in time.
A. G. Shypula, A. Madaan, Y. Zeng, U. Alon, J. R. Gardner, Y. Yang, M. Hashemi, G. Neubig, P. Ranganathan, O. Bastani, and A. Yazdanbakhsh, “Learning performance-improving code edits,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=ix7rLVHXyY
2024
Closest in time.
J. Chen, X. Hu, Z. Li, C. Gao, X. Xia, and D. Lo, “Code search is all you need? improving code suggestions with code search,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, ICSE 2024, Lisbon, Portugal, April 14-20, 2024 , 2024, pp. 73:1–73:13
2024
Closest in time.
“Models - Hugging Face,” https://huggingface.co/models , 2024, last accessed Mar. 2024
2024
Closest in time.
“API Reference - OpenAI API,” https://platform.openai.com/docs/api-reference , 2024, last accessed Mar. 2024
2024
Closest in time.
“openai_humaneval · Datasets at Hugging Face,” https://huggingface.co/datasets/openai_humaneval , 2024, last accessed Mar. 2024
2024
Closest in time.
“FudanSELab/ClassEval · Datasets at Hugging Face,” https://huggingface.co/datasets/FudanSELab/ClassEval , 2024, last accessed Mar. 2024
2024
Closest in time.
“Replication package,” https://figshare.com/s/e5de95bd79ab5ddea76c , 2024, replication package
2024
Closest in time.
2024
Closest in time.
Y. Peng, C. Gao, Z. Li, B. Gao, D. Lo, Q. Zhang, and M. Lyu, “Static inference meets deep learning: a hybrid type inference approach for python,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 2019–2030
2030
Closest in time.