Fetching the paper…
Reading the bibliography…
In this work, we make the first attempt to evaluate LLMs in a more challenging code generation scenario, i.e.
B. Meyer, “Applying "design by contract",” Computer , vol. 25, no. 10, pp. 40–51, 1992. [Online]. Available: https://doi.org/10.1109/2.161279
1992
Earlier work this paper cites.
T. Bhat and N. Nagappan, “Evaluating the efficacy of test-driven development: industrial case studies,” in 2006 International Symposium on Empirical Software Engineering (ISESE 2006), September 21-22, 2006, Rio de Janeiro, Brazil , G. H. Travassos, J. C. Maldonado, and C. Wohlin, Eds. ACM, 2006, pp. 356–363. [Online]. Available: https://doi.org/10.1145/1159733.1159787
2006
Earlier work this paper cites.
S. Chen, R. Varma, A. Sandryhaila, and J. Kovacevic, “Discrete signal processing on graphs: Sampling theory,” IEEE Trans. Signal Process. , vol. 63, no. 24, pp. 6510–6523, 2015. [Online]. Available: https://doi.org/10.1109/TSP.2015.2469645
2015
Earlier work this paper cites.
K. Srinath, “Python–the fastest growing programming language,” International Research Journal of Engineering and Technology , vol. 4, no. 12, pp. 354–357, 2017
2017
Earlier work this paper cites.
S. Iyer, I. Konstas, A. Cheung, and L. Zettlemoyer, “Mapping language to code in programmatic context,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds. Association for Computational Linguistics, 2018, pp. 1643–1652. [Online]. Available: https://doi.org/10.18653/v1/d18-1192
2018
Earlier work this paper cites.
P. Yin, B. Deng, E. Chen, B. Vasilescu, and G. Neubig, “Learning to mine aligned code and natural language pairs from stack overflow,” in Proceedings of the 15th International Conference on Mining Software Repositories, MSR 2018, Gothenburg, Sweden, May 28-29, 2018 , A. Zaidman, Y. Kamei, and E. Hill, Eds. ACM, 2018, pp. 476–486. [Online]. Available: https://doi.org/10.1145/3196398.3196408
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net, 2020. [Online]. Available: https://openreview.net/forum?id=rygGQyrFvH
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt, “Measuring coding challenge competence with APPS,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual , J. Vanschoren and S. Yeung, Eds., 2021. [Online]. Available: https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/c24cd76e1ce41366a4bbe8a49b02a028-Abstract-round2.html
2021
Earlier work this paper cites.
(2021) Dense-6.7b. [Online]. Available: https://huggingface.co/KoboldAI/fairseq-dense-6.7B-Shinen
2021
Earlier work this paper cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “GLM: general language model pretraining with autoregressive blank infilling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Association for Computational Linguistics, 2022, pp. 320–335. [Online]. Available: https://doi.org/10.18653/v1/2022.acl-long.26
2022
Earlier work this paper cites.
F. F. Xu, U. Alon, G. Neubig, and V. J. Hellendoorn, “A systematic evaluation of large language models of code,” in MAPS@PLDI 2022: 6th ACM SIGPLAN International Symposium on Machine Programming, San Diego, CA, USA, 13 June 2022 , S. Chaudhuri and C. Sutton, Eds. ACM, 2022, pp. 1–10. [Online]. Available: https://doi.org/10.1145/3520312.3534862
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
V. Vikram, C. Lemieux, and R. Padhye, “Can large language models write good property-based tests?” 2023
2023
Closest in time.
B. Shen, J. Zhang, T. Chen, D. Zan, B. Geng, A. Fu, M. Zeng, A. Yu, J. Ji, J. Zhao, Y. Guo, and Q. Wang, “Pangu-coder2: Boosting large language models for code with ranking feedback,” 2023
2023
Closest in time.
D. Zan, B. Chen, F. Zhang, D. Lu, B. Wu, B. Guan, W. Yongji, and J.-G. Lou, “Large language models meet NL2Code: A survey,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 7443–7464. [Online]. Available: https://aclanthology.org/2023.acl-long.411
2023
Closest in time.
(2023) Instruct-starcoder. [Online]. Available: https://huggingface.co/GeorgiaTechResearchInstitute/starcoder-gpteacher-code-instruct
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
S. Kang, J. Yoon, and S. Yoo, “Large language models are few-shot testers: Exploring llm-based general bug reproduction,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) , 2023, pp. 2312–2323
2023
Cited alongside, same era.
S. Kang, B. Chen, S. Yoo, and J.-G. Lou, “Explainable automated debugging via large language model-driven scientific debugging,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
(2023) Instruct-codegen. [Online]. Available: https://huggingface.co/sahil2801/instruct-codegen-16B
2023
Cited alongside, same era.
2023
Closest in time.
B. Athiwaratkun, S. K. Gouda, Z. Wang, X. Li, Y. Tian, M. Tan, W. U. Ahmad, S. Wang, Q. Sun, M. Shang, S. K. Gonugondla, H. Ding, V. Kumar, N. Fulton, A. Farahani, S. Jain, R. Giaquinto, H. Qian, M. K. Ramanathan, and R. Nallapati, “Multi-lingual evaluation of code generation models,” in The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023. [Online]. Available: https://openreview.net/pdf?id=Bo7eeXm6An8
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia, W. L. Tam, Z. Ma, Y. Xue, J. Zhai, W. Chen, Z. Liu, P. Zhang, Y. Dong, and J. Tang, “GLM-130B: an open bilingual pre-trained model,” in The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023. [Online]. Available: https://openreview.net/pdf?id=-Aw0rrrPUF
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Chervenak, H. Lieman, M. Blanco-Breindel, and S. Jindal, “The promise and peril of using a large language model to obtain clinical information: Chatgpt performs strongly as a fertility counseling tool with limitations,” Fertility and Sterility , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.