Fetching the paper…
Reading the bibliography…
The success of language models in code assistance has spurred the proposal of repository-level code completion as a means to enhance prediction accuracy, utilizing the context from the entire codebase.
Mining source code repositories at massive scale using language modeling
Allamanis, M. and Sutton, C · 2013
Earlier work this paper cites.
Modeling and discovering vulnerabilities with code property graphs
Yamaguchi, F., Golde, N., Arp, D., and Rieck, K · 2014
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
CodeBERT: A pre-trained model for programming and natural languages
Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., Jiang, D., and Zhou, M · 2020
Earlier work this paper cites.
Big code != big vocabulary: open-vocabulary models for source code
Karampatsis, R.-M., Babii, H., Robbes, R., Sutton, C., and Janes, A · 2020
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Efficient training of language models to fill in the middle
Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., and Chen, M · 2022
Earlier work this paper cites.
Pangu-coder: Program synthesis with function-level language modeling
Christopoulou, F., Lampouras, G., Gritta, M., Zhang, G., Guo, Y., Li, Z., Zhang, Q., Xiao, M., Shen, B., Li, L., Yu, H., Yan, L., Zhou, P., Wang, X., Ma, Y., Iacobacci, I., Wang, Y., Liang, G., Wei, J., Jiang, X., Wang, Q., and Liu, Q · 2022
Earlier work this paper cites.
Cocomic: Code completion by jointly modeling in-file and cross-file context
Ding, Y., Wang, Z., Ahmad, W. U., Ramanathan, M. K., Nallapati, R., Bhatia, P., Roth, D., and Xiang, B · 2022
Earlier work this paper cites.
UniXcoder: Unified cross-modal pre-training for code representation
Guo, D., Lu, S., Duan, N., Wang, Y., Zhou, M., and Yin, J · 2022
Earlier work this paper cites.
ReACC: A retrieval-augmented code completion framework
Lu, S., Duan, N., Han, H., Guo, D., Hwang, S.-w., and Svyatkovskiy, A · 2022
Cited alongside, same era.
Santacoder: don’t reach for the stars!
Allal, L. B., Li, R., Kocetkov, D., Mou, C., Akiki, C., Ferrandis, C. M., Muennighoff, N., Mishra, M., Gu, A., Dey, M., Umapathi, L. K., Anderson, C. J., Zi, Y., Lamy-Poirier, J., Schoelkopf, H., Troshin, S., Abulkhanov, D., Romero, M., Lappert, M., Toni, F. D., del Río, B. G., Liu, Q., Bose, S., Bhattacharyya, U., Zhuo, T. Y., Yu, I., Villegas, P., Zocca, M., Mangrulkar, S., Lansky, D., Nguyen, H., Contractor, D., Villa, L., Li, J., Bahdanau, D., Jernite, Y., Hughes, S., Fried, D., Guha, A., de Vries, H., and von Werra, L · 2023
Cited alongside, same era.
Codeplan: Repository-level coding using LLMs and planning
Bairi, R., Sonwane, A., Kanade, A., C, V. D., Iyer, A., Parthasarathy, S., Rajamani, S., Ashok, B., and Shet, S · 2023
Cited alongside, same era.
Grounded copilot: How programmers interact with code-generating models
Barke, S., James, M. B., and Polikarpova, N · 2023
Cited alongside, same era.
Codefuse-13b: A pretrained multi-lingual code large language model
Codegen2: Lessons for training llms on programming and natural languages
Nijkamp, E., Hayashi, H., Xiong, C., Savarese, S., and Zhou, Y · 2023
Later among the works it cites.
Better context makes better code language models: A case study on function call argument completion
Pei, H., Zhao, J., Lausen, L., Zha, S., and Karypis, G · 2023
Later among the works it cites.
Code llama: Open foundation models for code, 2023
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2023
Later among the works it cites.
How practitioners expect code completion?
Wang, C., Hu, J., Gao, C., Jin, Y., Xie, T., Huang, H., Lei, Z., and Deng, Y · 2023
Later among the works it cites.
A systematic literature review on source code similarity measurement and clone detection: techniques, applications, and challenges, 2023
Zakeri-Nasrabadi, M., Parsa, S., Ramezani, M., Roy, C., and Ekhtiarzadeh, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Di, P., Li, J., Yu, H., Jiang, W., Cai, W., Cao, Y., Chen, C., Chen, D., Chen, H., Chen, L., Fan, G., Gong, J., Gong, Z., Hu, W., Guo, T., Lei, Z., Li, T., Li, Z., Liang, M., Liao, C., Liu, B., Liu, J., Liu, Z., Lu, S., Shen, M., Wang, G., Wang, H., Wang, Z., Xu, Z., Yang, J., Ye, Q., Zhang, G., Zhang, Y., Zhao, Z., Zheng, X., Zhou, H., Zhu, L., and Zhu, X · 2023
Cited alongside, same era.
Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion, 2023
Ding, Y., Wang, Z., Ahmad, W. U., Ding, H., Tan, M., Jain, N., Ramanathan, M. K., Nallapati, R., Bhatia, P., Roth, D., and Xiang, B · 2023
Cited alongside, same era.
Incoder: A generative model for code infilling and synthesis
Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, S., Zettlemoyer, L., and Lewis, M · 2023
Cited alongside, same era.
Grace: Language models meet code edits
Gupta, P., Khare, A., Bajpai, Y., Chakraborty, S., Gulwani, S., Kanade, A., Radhakrishna, A., Soares, G., and Tiwari, A · 2023
Cited alongside, same era.
Starcoder: may the source be with you!, 2023
Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., Liu, Q., Zheltonozhskii, E., Zhuo, T. Y., Wang, T., Dehaene, O., Davaadorj, M., Lamy-Poirier, J., Monteiro, J., Shliazhko, O., Gontier, N., Meade, N., Zebaze, A., Yee, M.-H., Umapathi, L. K., Zhu, J., Lipkin, B., Oblokulov, M., Wang, Z., Murthy, R., Stillerman, J., Patel, S. S., Abulkhanov, D., Zocca, M., Dey, M., Zhang, Z., Fahmy, N., Bhattacharyya, U., Yu, W., Singh, S., Luccioni, S., Villegas, P., Kunakov, M., Zhdanov, F., Romero, M., Lee, T., Timor, N., Ding, J., Schlesinger, C., Schoelkopf, H., Ebert, J., Dao, T., Mishra, M., Gu, A., Robinson, J., Anderson, C. J., Dolan-Gavitt, B., Contractor, D., Reddy, S., Fried, D., Bahdanau, D., Jernite, Y., Ferrandis, C. M., Hughes, S., Wolf, T., Guha, A., von Werra, L., and de Vries, H · 2023
Cited alongside, same era.
Repobench: Benchmarking repository-level code auto-completion systems, 2023
Liu, T., Xu, C., and McAuley, J · 2023
Cited alongside, same era.
GitHub - davidhalter/jedi: Awesome autocompletion, static analysis and refactoring library for python — github.com
Dave Halter
Cited in the paper.
Repofusion: Training code models to understand your repository, 2023a
Shrivastava, D., Kocetkov, D., de Vries, H., Bahdanau, D., and Scholak, T
Cited in the paper.
Later among the works it cites.
Private-library-oriented code generation with large language models, 2023
Zan, D., Chen, B., Gong, Y., Cao, J., Zhang, F., Wu, B., Guan, B., Yin, Y., and Wang, Y · 2023
Later among the works it cites.
RepoCoder: Repository-level code completion through iterative retrieval and generation
Zhang, F., Chen, B., Zhang, Y., Keung, J., Liu, J., Zan, D., Mao, Y., Lou, J.-G., and Chen, W · 2023
Later among the works it cites.
Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x
Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Wang, Z., Shen, L., Wang, A., Li, Y., Su, T., Yang, Z., and Tang, J · 2023
Later among the works it cites.
Deepseek-coder: When the large language model meets programming – the rise of code intelligence, 2024
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y. K., Luo, F., Xiong, Y., and Liang, W · 2024
Closest in time.