Fetching the paper…
Reading the bibliography…
Large language models have made substantial progress in addressing diverse code-related tasks.
T. Breaux and A. Antón, “Analyzing regulatory rules for privacy and security requirements,” IEEE Transactions on Software Engineering , vol. 34, no. 1, pp. 5–20, 2008
2008
Earlier work this paper cites.
J. Bloch, Effective java (the java series) . Prentice Hall PTR, 2008
2008
Earlier work this paper cites.
R. Vallée-Rai, P. Co, E. Gagnon, L. Hendren, P. Lam, and V. Sundaresan, “Soot: A java bytecode optimization framework,” in CASCON First Decade High Impact Papers . IBM, 2010, p. 214–224
2010
Earlier work this paper cites.
C. McMillan, M. Grechanik, D. Poshyvanyk, Q. Xie, and C. Fu, “Portfolio: finding relevant functions and their usage,” in International Conference on Software Engineering , 2011, pp. 111–120
2011
Earlier work this paper cites.
R. P. L. Buse and W. Weimer, “Synthesizing api usage examples,” in International Conference on Software Engineering . IEEE Press, 2012, p. 782–792
2012
Earlier work this paper cites.
T. Gvero and V. Kuncak, “Synthesizing java expressions from free-form queries,” SIGPLAN Not. , vol. 50, no. 10, p. 416–432, 2015
2015
Earlier work this paper cites.
M. Raghothaman, Y. Wei, and Y. Hamadi, “Swim: synthesizing what i mean: code search and idiomatic snippet synthesis,” in International Conference on Software Engineering . ACM, 2016, p. 357–367
2016
Earlier work this paper cites.
M. M. Rahman, C. K. Roy, and D. Lo, “Rack: Automatic api recommendation using crowdsourced knowledge,” in International Conference on Software Analysis, Evolution, and Reengineering (SANER) , vol. 1. IEEE, 2016, pp. 349–359
2016
Earlier work this paper cites.
X. Gu, H. Zhang, D. Zhang, and S. Kim, “Deep api learning,” in Foundations of Software Engineering . ACM, 2016, p. 631–642
2016
Earlier work this paper cites.
P. Yin and G. Neubig, “A syntactic neural model for general-purpose code generation,” in 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , R. Barzilay and M.-Y. Kan, Eds. ACL, 2017, pp. 440–450
2017
Earlier work this paper cites.
Q. Huang, X. Xia, Z. Xing, D. Lo, and X. Wang, “Api method recommendation without worrying about the task-api knowledge gap,” in International Conference on Automated Software Engineering . ACM, 2018, p. 293–304
2018
Earlier work this paper cites.
M. M. Rahman and C. Roy, “Nlp2api: Query reformulation for code search using crowdsourced knowledge and extra-large data analytics,” in International Conference on Software Maintenance and Evolution . IEEE, 2018, pp. 714–714
2018
Earlier work this paper cites.
T. Yu, R. Zhang, K. Yang, M. Yasunaga, D. Wang, Z. Li, J. Ma, I. Li, Q. Yao, S. Roman, Z. Zhang, and D. Radev, “Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task,” in Empirical Methods in Natural Language Processing . ACL, 2018, pp. 3911–3921
2018
Earlier work this paper cites.
B. Dhingra, M. Faruqui, A. Parikh, M.-W. Chang, D. Das, and W. Cohen, “Handling divergent reference texts when evaluating table-to-text generation,” in Annual Meeting of the Association for Computational Linguistics . ACL, 2019, pp. 4884–4895
2019
Earlier work this paper cites.
M. Liu, X. Peng, A. Marcus, Z. Xing, W. Xie, S. Xing, and Y. Liu, “Generating query-specific class api summaries,” in Foundations of Software Engineering . ACM, 2019, p. 120–130
2019
Earlier work this paper cites.
T. H. M. Le, H. Chen, and M. A. Babar, “Deep learning for source code modeling and generation: Models, applications, and challenges,” ACM Comput. Surv. , vol. 53, no. 3, 2020
2020
Earlier work this paper cites.
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020 . ACL, 2020, pp. 1536–1547
2020
Earlier work this paper cites.
H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, and R. Karri, “An empirical cybersecurity evaluation of github copilot’s code contributions,” arXiv , pp. arXiv–2108, 2021
2021
Earlier work this paper cites.
N. Alhirabi, O. Rana, and C. Perera, “Security and privacy requirements for the internet of things: A survey,” ACM Transactions on Internet of Things , vol. 2, no. 1, pp. 1–37, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” in ACM Conference on Fairness, Accountability, and Transparency . ACM, 2021, p. 610–623
2021
Earlier work this paper cites.
Y. Wang, W. Wang, S. Joty, and S. C. Hoi, “CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation,” in Empirical Methods in Natural Language Processing . ACL, 2021, pp. 8696–8708
2021
Earlier work this paper cites.
D. Guo, S. Ren, S. Lu, Z. Feng, D. Tang, S. Liu, L. Zhou, N. Duan, A. Svyatkovskiy, S. Fu, M. Tufano, S. K. Deng, C. B. Clement, D. Drain, N. Sundaresan, J. Yin, D. Jiang, and M. Zhou, “Graphcodebert: Pre-training code representations with data flow,” in International Conference on Learning Representations . OpenReview.net, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
P. Jain, A. Jain, T. Zhang, P. Abbeel, J. Gonzalez, and I. Stoica, “Contrastive code representation learning,” in Empirical Methods in Natural Language Processing . ACL, 2021, pp. 5954–5971
2021
Earlier work this paper cites.
L. Phan, H. Tran, D. Le, H. Nguyen, J. Annibal, A. Peltekian, and Y. Ye, “CoTexT: Multi-task learning with code-text transformer,” in Workshop on Natural Language Processing for Programming , 2021, pp. 40–47
2021
Earlier work this paper cites.
W. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang, “Unified pre-training for program understanding and generation,” in Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . ACL, 2021, pp. 2655–2668
2021
Earlier work this paper cites.
N. Dziri, A. Madotto, O. Zaïane, and A. J. Bose, “Neural path hunter: Reducing hallucination in dialogue systems via path grounding,” in Empirical Methods in Natural Language Processing . ACL, 2021, pp. 2197–2214
2021
Earlier work this paper cites.
2022
Cited alongside, same era.
N. Nguyen and S. Nadi, “An empirical evaluation of GitHub Copilot’s code suggestions,” in ACM International Conference on Mining Software Repositories , 2022, pp. 1–5
2022
Cited alongside, same era.
P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,” in Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems . ACM, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
OpenAI, “Chat completions,” https://platform.openai.com/docs/guides/chat , 2023, accessed: December 6, 2023
2023
Later among the works it cites.
F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M. Yee, Y. Zi, C. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda, “Multipl-e: A scalable and polyglot approach to benchmarking neural code generation,” IEEE Transactions on Software Engineering , vol. 49, no. 07, pp. 3675–3691, 2023
2023
Later among the works it cites.
Q. Zheng, X. Xia, X. Zou, Y. Dong, S. Wang, Y. Xue, L. Shen, Z. Wang, A. Wang, Y. Li, T. Su, Z. Yang, and J. Tang, “Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x,” in SIGKDD Conference on Knowledge Discovery and Data Mining . ACM, 2023, p. 5673–5684
2023
Later among the works it cites.
Y. Peng, S. Li, W. Gu, Y. Li, W. Wang, C. Gao, and M. R. Lyu, “Revisiting, benchmarking and exploring api recommendation: How far are we?” Trans. Softw. Eng. , vol. 49, no. 4, p. 1876–1897, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Imai, “Is github copilot a substitute for human pair-programming? an empirical study,” in 44th International Conference on Software Engineering (ICSE-Companion) , 2022, pp. 319–321
2022
Cited alongside, same era.
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys , 2022
2022
Cited alongside, same era.
F. F. Xu, U. Alon, G. Neubig, and V. J. Hellendoorn, “A systematic evaluation of large language models of code,” in International Symposium on Machine Programming . ACM, 2022, p. 1–10
2022
Cited alongside, same era.
“Codex model,” https://beta.openai.com/docs/models/codex-series-private-beta , 2022
2022
Cited alongside, same era.
A. Ziegler, E. Kalliamvakou, X. A. Li, A. Rice, D. Rifkin, S. Simister, G. Sittampalam, and E. Aftandilian, “Productivity assessment of neural code completion,” in International Symposium on Machine Programming . ACM, 2022, p. 21–29
2022
Cited alongside, same era.
T. Liu, Y. Zhang, C. Brockett, Y. Mao, Z. Sui, W. Chen, and B. Dolan, “A token-level reference-free hallucination detection benchmark for free-form text generation,” in Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . ACL, 2022, pp. 6723–6737
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. M. Mcnutt, C. Wang, R. A. Deline, and S. M. Drucker, “On the design of ai-powered code assistants for notebooks,” in CHI Conference on Human Factors in Computing Systems . ACM, 2023
2023
Later among the works it cites.
“Tree-sitter,” https://tree-sitter.github.io/tree-sitter/ , 2023
2023
Later among the works it cites.
C. Deng, Y. Zhao, X. Tang, M. Gerstein, and A. Cohan, “Benchmark probing: Investigating data leakage in large language models,” in NeurIPS 2023 Workshop on Backdoors in Deep Learning-The Good, the Bad, and the Ugly , 2023
2023
Later among the works it cites.
T. Ahmed and P. Devanbu, “Few-shot training llms for project-specific code-summarization,” in International Conference on Automated Software Engineering . ACM, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Dai, Z. Liu, Z. Ji, D. Su, and P. Fung, “Plausible may not be faithful: Probing object hallucination in vision-language pre-training,” in 17th Conference of the European Chapter of the Association for Computational Linguistics . ACL, 2023, pp. 2136–2148
2023
Later among the works it cites.
P. Manakul, A. Liusie, and M. Gales, “SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models,” in Empirical Methods in Natural Language Processing . ACL, 2023, pp. 9004–9017
2023
Later among the works it cites.
OpenAI, “Gpt-4 technical report,” 2023
2023
Later among the works it cites.
D. Shrivastava, H. Larochelle, and D. Tarlow, “Repository-level prompt generation for large language models of code,” in International Conference on Machine Learning . PMLR, 2023, pp. 31 693–31 715
2023
Later among the works it cites.
M. Liu, T. Yang, Y. Lou, X. Du, Y. Wang, and X. Peng, “Codegen4libs: A two-stage approach for library-oriented code generation,” in International Conference on Automated Software Engineering . IEEE, 2023, pp. 434–445
2023
Later among the works it cites.
S. Zhou, U. Alon, F. F. Xu, Z. Jiang, and G. Neubig, “Docprompting: Generating code by retrieving the docs,” in International Conference on Learning Representations , 2023
2023
Later among the works it cites.
H. Pei, J. Zhao, L. Lausen, S. Zha, and G. Karypis, “Better context makes better code language models: A case study on function call argument completion,” in AAAI Conference on Artificial Intelligence , vol. 37, no. 4, 2023, pp. 5230–5238
2023
Later among the works it cites.
Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, W.-t. Yih, D. Fried, S. Wang, and T. Yu, “Ds-1000: A natural and reliable benchmark for data science code generation,” in International Conference on Machine Learning . PMLR, 2023, pp. 18 319–18 345
2023
Later among the works it cites.
A. Ziegler, E. Kalliamvakou, X. A. Li, A. Rice, D. Rifkin, S. Simister, G. Sittampalam, and E. Aftandilian, “Measuring github copilot’s impact on productivity,” Communications of the ACM , vol. 67, no. 3, pp. 54–63, 2024
2024
Closest in time.
2024
Closest in time.
A. Al-Kaswan, M. Izadi, and A. V. Deursen, “Traces of memorisation in large language models for code,” in International Conference on Software Engineering . IEEE Computer Society, 2024, pp. 862–862
2024
Closest in time.
Z. Yang, Z. Zhao, C. Wang, J. Shi, D. Kim, D. Han, and D. Lo, “Unveiling memorization in code models,” in International Conference on Software Engineering . IEEE Computer Society, 2024, pp. 856–856
2024
Closest in time.
X. Du, M. Liu, K. Wang, H. Wang, J. Liu, Y. Chen, J. Feng, C. Sha, X. Peng, and Y. Lou, “Evaluating large language models in class-level code generation,” in International Conference on Software Engineering . IEEE Computer Society, 2024, pp. 865–865
2024
Closest in time.
2024
Closest in time.
“Tiobe language popularity index,” https://www.tiobe.com/tiobe-index/ , 2024, accessed: 2024-03-15
2024
Closest in time.
2024
Closest in time.
H. Yu, B. Shen, D. Ran, J. Zhang, Q. Zhang, Y. Ma, G. Liang, Y. Li, Q. Wang, and T. Xie, “Codereval: A benchmark of pragmatic code generation with generative pre-trained models,” in International Conference on Software Engineering , 2024, pp. 1–12
2024
Closest in time.