Fetching the paper…
Reading the bibliography…
The rapid advancements in large language models (LLMs) have greatly expanded the potential for automated code-related tasks.
M. B. Miles and A. M. Huberman, Qualitative data analysis: An expanded sourcebook . sage, 1994
1994
Earlier work this paper cites.
S. V. Stehman, “Selecting and interpreting measures of thematic classification accuracy,” Remote sensing of Environment , vol. 62, no. 1, pp. 77–89, 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
B. Kitchenham and S. L. Pfleeger, “Principles of survey research: part 5: populations and samples,” ACM SIGSOFT Software Engineering Notes , vol. 27, no. 5, pp. 17–20, 2002
2002
Earlier work this paper cites.
2005
Earlier work this paper cites.
G. Guest, A. Bunce, and L. Johnson, “How many interviews are enough? an experiment with data saturation and variability,” Field methods , vol. 18, no. 1, pp. 59–82, 2006
2006
Earlier work this paper cites.
S. Robertson, H. Zaragoza et al. , “The probabilistic relevance framework: Bm25 and beyond,” Foundations and Trends® in Information Retrieval , vol. 3, no. 4, pp. 333–389, 2009
2009
Earlier work this paper cites.
S. Wang, T. Liu, and L. Tan, “Automatically learning semantic features for defect prediction,” in Proceedings of the 38th International Conference on Software Engineering , 2016, pp. 297–308
2016
Earlier work this paper cites.
M. White, M. Tufano, C. Vendome, and D. Poshyvanyk, “Deep learning code fragments for code clone detection,” in Proceedings of the 31st IEEE/ACM international conference on automated software engineering , 2016, pp. 87–98
2016
Earlier work this paper cites.
P. Yin and G. Neubig, “Tranx: A transition-based neural abstract syntax parser for semantic parsing and code generation,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing (Demo Track) , 2018
2018
Earlier work this paper cites.
M. Allamanis, M. Brockschmidt, and M. Khademi, “Learning to represent programs with graphs,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
X. Gu, H. Zhang, and S. Kim, “Deep code search,” in Proceedings of the 40th International Conference on Software Engineering , 2018, pp. 933–944
2018
Earlier work this paper cites.
X. Hu, G. Li, X. Xia, D. Lo, and Z. Jin, “Deep code comment generation,” in Proceedings of the 26th conference on program comprehension , 2018, pp. 200–210
2018
Earlier work this paper cites.
H. Liu, J. Jin, Z. Xu, Y. Zou, Y. Bu, and L. Zhang, “Deep learning based code smell detection,” IEEE transactions on Software Engineering , vol. 47, no. 9, pp. 1811–1837, 2019
2019
Earlier work this paper cites.
H. Husain, H.-H. Wu, T. Gazit, M. Allamanis, and M. Brockschmidt, “Codesearchnet challenge: Evaluating the state of semantic code search,” 2019
2019
Earlier work this paper cites.
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang et al. , “Codebert: A pre-trained model for programming and natural languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , 2020, pp. 1536–1547
2020
Earlier work this paper cites.
A. Kanade, P. Maniatis, G. Balakrishnan, and K. Shi, “Learning and evaluating contextual embedding of source code,” in International conference on machine learning . PMLR, 2020, pp. 5110–5121
2020
Earlier work this paper cites.
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh, “Autoprompt: Eliciting knowledge from language models with automatically generated prompts,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 4222–4235
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
X. Hu, G. Li, X. Xia, D. Lo, and Z. Jin, “Deep code comment generation with hybrid lexical and syntactical information,” Empirical Software Engineering , vol. 25, pp. 2179–2217, 2020
2020
Earlier work this paper cites.
J. Shin and J. Nam, “A survey of automatic code generation from natural language,” Journal of Information Processing Systems , vol. 17, no. 3, pp. 537–555, 2021
2021
Earlier work this paper cites.
W. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang, “Unified pre-training for program understanding and generation,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 2655–2668
2021
Earlier work this paper cites.
K. Hambardzumyan, H. Khachatrian, and J. May, “Warp: Word-level adversarial reprogramming,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 4921–4933
2021
Earlier work this paper cites.
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang, G. Li, L. Zhou, L. Shou, L. Zhou, M. Tufano, M. Gong, M. Zhou, N. Duan, N. Sundaresan, S. K. Deng, S. Fu, and S. Liu, “Codexglue: A machine learning benchmark dataset for code understanding and generation,” 2021
2021
Earlier work this paper cites.
C. et al., “Evaluating large language models trained on code,” 2021
2021
Earlier work this paper cites.
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton, “Program synthesis with large language models,” 2021
2021
Earlier work this paper cites.
L. Phan, H. Tran, D. Le, H. Nguyen, J. Annibal, A. Peltekian, and Y. Ye, “Cotext: Multi-task learning with code-text transformer,” 2021. [Online]. Available: http://dx.doi.org/10.18653/v1/2021.nlp4prog-1.5
2021
Earlier work this paper cites.
W. Qi, Y. Gong, Y. Yan, C. Xu, B. Yao, B. Zhou, B. Cheng, D. Jiang, J. Chen, R. Zhang, H. Li, and N. Duan, “ProphetNet-X: Large-scale pre-training models for English, Chinese, multi-lingual, dialog, and code generation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations , Online, Aug. 2021, pp. 232–239. [Online]. Available: https://aclanthology.org/2021.acl-demo.28
2021
Earlier work this paper cites.
Y. Wang, W. Wang, S. Joty, and S. C. Hoi, “Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 8696–8708
2021
Earlier work this paper cites.
W. Sun, C. Fang, Y. Chen, G. Tao, T. Han, and Q. Zhang, “Code search based on context-aware code translation,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 388–400
2022
Cited alongside, same era.
Y. Li, S. Wang, and T. N. Nguyen, “Dear: A novel deep learning-based approach for automated program repair,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 511–523
2022
Cited alongside, same era.
D. Guo, S. Lu, N. Duan, Y. Wang, M. Zhou, and J. Yin, “Unixcoder: Unified cross-modal pre-training for code representation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 7212–7225
2022
Cited alongside, same era.
T. Ahmed and P. Devanbu, “Multilingual training for software engineering,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 1443–1455
2022
Cited alongside, same era.
S. Ouyang, J. M. Zhang, M. Harman, and M. Wang, “Llm is like a box of chocolates: the non-determinism of chatgpt in code generation,” 2023
2023
Closest in time.
H. Li, Y. Hao, Y. Zhai, and Z. Qian, “Assisting static analysis with large language models: A chatgpt experiment,” in Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2023, pp. 2107–2111
2023
Closest in time.
2023
Closest in time.
S. Tipirneni, M. Zhu, and C. K. Reddy, “Structcoder: Structure-aware transformer for code generation,” 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” 2022
2022
Cited alongside, same era.
A. C. et al., “Palm: Scaling language modeling with pathways,” 2022
2022
Cited alongside, same era.
J. Y. Khan and G. Uddin, “Automatic code documentation generation using gpt-3,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering , 2022, pp. 1–6
2022
Cited alongside, same era.
J. Liu, D. Shen, Y. Zhang, B. Dolan, L. Carin, and W. Chen, “What makes good in-context examples for gpt-3?” in Deep Learning Inside Out: 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, DeeLIO 2022 . Association for Computational Linguistics (ACL), 2022, pp. 100–114
2022
Cited alongside, same era.
Y. He, S. Zheng, Y. Tay, J. Gupta, Y. Du, V. Aribandi, Z. Zhao, Y. Li, Z. Chen, D. Metzler et al. , “Hyperprompt: Prompt-based task-conditioning of transformers,” in International Conference on Machine Learning . PMLR, 2022, pp. 8678–8690
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 824–24 837, 2022
2022
Cited alongside, same era.
C. Wang, Y. Yang, C. Gao, Y. Peng, H. Zhang, and M. R. Lyu, “No more fine-tuning? an experimental evaluation of prompt tuning in code intelligence,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2022, pp. 382–394
2022
Cited alongside, same era.
T. Ahmed and P. Devanbu, “Few-shot training llms for project-specific code-summarization,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering , 2022, pp. 1–5
2022
Cited alongside, same era.
2023
Closest in time.
Z. Yuan, J. Liu, Q. Zi, M. Liu, X. Peng, and Y. Lou, “Evaluating instruction-tuned large language models on code comprehension and generation,” 2023
2023
Closest in time.
J. Yang, H. Jin, R. Tang, X. Han, Q. Feng, H. Jiang, S. Zhong, B. Yin, and X. Hu, “Harnessing the power of llms in practice: A survey on chatgpt and beyond,” ACM Transactions on Knowledge Discovery from Data , 2023
2023
Closest in time.
Y. Chen, Q. Fu, Y. Yuan, Z. Wen, G. Fan, D. Liu, D. Zhang, Z. Li, and Y. Xiao, “Hallucination detection: Robustly discerning reliable answers in large language models,” in Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , 2023, pp. 245–255
2023
Closest in time.
N. M. Guerreiro, D. M. Alves, J. Waldendorf, B. Haddow, A. Birch, P. Colombo, and A. F. Martins, “Hallucinations in large multilingual translation models,” Transactions of the Association for Computational Linguistics , vol. 11, pp. 1500–1517, 2023
2023
Closest in time.
D. Trautmann, “Large language model prompt chaining for long legal document classification,” SwissText’23: The 8th edition of the Swiss Text Analytics Conference – Generative AI & LLM, June 12–14, 2023, Neuchâtel, Switzerland , 2023
2023
Closest in time.
S. MacNeil, A. Tran, A. Hellas, J. Kim, S. Sarsa, P. Denny, S. Bernstein, and J. Leinonen, “Experiences from using code explanations generated by large language models in a web software development e-book,” in Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 , 2023, pp. 931–937
2023
Closest in time.
J. Shin, S. Hashtroudi, H. Hemmati, and S. Wang, “Domain adaptation for code model-based unit test case generation,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis , 2024, pp. 1211–1222
2024
Closest in time.
2024
Closest in time.
J. Shin, H. Hemmati, M. Wei, and S. Wang, “Assessing evaluation metrics for neural test oracle generation,” IEEE Transactions on Software Engineering , 2024
2024
Closest in time.
M. Geng, S. Wang, D. Dong, H. Wang, G. Li, Z. Jin, X. Mao, and X. Liao, “Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning,” in 2024 IEEE/ACM 46th International Conference on Software Engineering (ICSE) , 2024
2024
Closest in time.
F. Trad and A. Chehab, “Prompt engineering or fine-tuning? a case study on phishing detection with large language models,” Machine Learning and Knowledge Extraction , vol. 6, no. 1, pp. 367–384, 2024
2024
Closest in time.
T. Ahmed, K. S. Pai, P. Devanbu, and E. Barr, “Automatic semantic augmentation of language model prompts (for code summarization),” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , 2024, pp. 1–13
2024
Closest in time.
F. Mu, L. Shi, S. Wang, Z. Yu, B. Zhang, C. Wang, S. Liu, and Q. Wang, “Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification,” Proceedings of the ACM on Software Engineering , vol. 1, no. FSE, pp. 2332–2354, 2024
2024
Closest in time.
Y. Dong, X. Jiang, Z. Jin, and G. Li, “Self-collaboration code generation via chatgpt,” ACM Transactions on Software Engineering and Methodology , vol. 33, no. 7, pp. 1–38, 2024
2024
Closest in time.
Y. Luo, R. Yu, F. Zhang, L. Liang, and Y. Xiong, “Bridging gaps in llm code translation: Reducing errors with call graphs and bridged debuggers,” in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering , 2024, pp. 2448–2449
2024
Closest in time.
R. Pan, A. R. Ibrahimzada, R. Krishna, D. Sankar, L. P. Wassi, M. Merler, B. Sobolev, R. Pavuluri, S. Sinha, and R. Jabbarvand, “Lost in translation: A study of bugs introduced by large language models while translating code,” in 2024 IEEE/ACM 46th International Conference on Software Engineering (ICSE) , 2024
2024
Closest in time.
Z. Yang, F. Liu, Z. Yu, J. W. Keung, J. Li, S. Liu, Y. Hong, X. Ma, Z. Jin, and G. Li, “Exploring and unleashing the power of large language models in automated code translation,” Proceedings of the ACM on Software Engineering , vol. 1, no. FSE, pp. 1585–1608, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Alshahwan, J. Chheda, A. Finegenova, B. Gokkaya, M. Harman, I. Harper, A. Marginean, S. Sengupta, and E. Wang, “Automated unit test improvement using large language models at meta,” 2024
2024
Closest in time.
E. A. AlOmar, A. Venkatakrishnan, M. W. Mkaouer, C. D. Newman, and A. Ouni, “How to refactor this code? an exploratory study on developer-chatgpt refactoring conversations,” 2024
2024
Closest in time.
J. Chen, H. Lin, X. Han, and L. Sun, “Benchmarking large language models in retrieval-augmented generation,” 2024
2024
Closest in time.
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.