Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are becoming an integrated part of software development.
P. J. Rousseeuw, “Least median of squares regression,” Journal of the American Statistical Association , vol. 79, no. 388, pp. 871–880, 1984
1984
Earlier work this paper cites.
F. T. Liu, K. M. Ting, and Z.-H. Zhou, “Isolation forest,” in Eighth IEEE International Conference on Data Mining , 2008, pp. 413–422
2008
Earlier work this paper cites.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research , vol. 12, pp. 2825–2830, 2011
2011
Earlier work this paper cites.
J. Svajlenko, J. F. Islam, I. Keivanloo, C. K. Roy, and M. M. Mia, “Towards a big data curated benchmark of inter-project code clones,” in Proceedings of the 2014 IEEE International Conference on Software Maintenance and Evolution , ser. ICSME ’14. USA: IEEE Computer Society, 2014, p. 476–480
2014
Earlier work this paper cites.
Y. H. Dovoedo and S. Chakraborti, “Boxplot-based outlier detection for the location-scale family,” Communications in Statistics—Simulation and Computation , 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems, Part of Advances in Neural Information Processing Systems, Volume 30 , ser. NIPS 2017. Red Hook, NY, USA: Curran Associates Inc., 2017, pp. 5998–6008
2017
Earlier work this paper cites.
B. Tran, J. Li, and A. Madry, “Spectral signatures in backdoor attacks,” Advances in neural information processing systems (NeurIPS) , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
K. Liu, B. Dolan-Gavitt, and S. Garg, “Fine-pruning: Defending against backdooring attacks on deep neural networks,” in International symposium on research in attacks, intrusions, and defenses . Springer, 2018, pp. 273–294
2018
Earlier work this paper cites.
M. Allamanis, “The adverse effects of code duplication in machine learning models of code,” in Proceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software (Onward!) , 2019, pp. 143–153
2019
Earlier work this paper cites.
Y. Zhou, S. Liu, J. Siow, X. Du, and Y. Liu, Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks . Red Hook, NY, USA: Curran Associates Inc., 2019
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 . Minneapolis, Minnesota: Association for Computational Linguistics, 2019, pp. 4171–4186
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
P. Bielik and M. Vechev, “Adversarial robustness for code,” in International Conference on Machine Learning . PMLR, 2020, pp. 896–907
2020
Earlier work this paper cites.
W. Wang, G. Li, B. Ma, X. Xia, and Z. Jin, “Detecting code clones with graph neural network and flow-augmented abstract syntax tree,” in 2020 IEEE 27th International Conference on Software Analysis, Evolution and Reengineering (SANER) . Los Alamitos, CA, USA: IEEE Computer Society, feb 2020, pp. 261–271
2020
Earlier work this paper cites.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A robustly optimized BERT pretraining approach,” International Conference on Learning Representations (ICLR), OpenReview.net , 2020
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, Jul. 2020, pp. 7871–7880
2020
Earlier work this paper cites.
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of the Association for Computational Linguistics: EMNLP . Online: Association for Computational Linguistics, 2020, pp. 1536–1547
2020
Earlier work this paper cites.
Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis, and L. Zettlemoyer, “Multilingual denoising pre-training for neural machine translation,” Transactions of the Association for Computational Linguistics , vol. 8, pp. 726–742, 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
M. R. I. Rabin, N. D. Bui, K. Wang, Y. Yu, L. Jiang, and M. A. Alipour, “On the generalizability of neural program models with respect to semantic-preserving program transformations,” Information and Software Technology (IST) , vol. 135(106552), pp. 1–13, 2021
2021
Cited alongside, same era.
S. Suneja, Y. Zheng, Y. Zhuang, J. A. Laredo, and A. Morari, “Probing model signal-awareness via prediction-preserving input minimization,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE) , 2021, pp. 945–955
2021
Cited alongside, same era.
M. R. I. Rabin, V. J. Hellendoorn, and M. A. Alipour, “Understanding neural code intelligence through program simplification,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE) , 2021, pp. 441–452
2021
Cited alongside, same era.
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong, “Codegen: An open large language model for code with multi-turn program synthesis,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=iaYcJKpY2B_
2023
Closest in time.
J. Zhang, S. Panthaplackel, P. Nie, J. J. Li, and M. Gligoric, “Coditt5: Pretraining for source code and natural language editing,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering , ser. ASE ’22. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3551349.3556955
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Schuster, C. Song, E. Tromer, and V. Shmatikov, “You autocomplete me: Poisoning vulnerabilities in neural code completion,” in 30th USENIX Security Symposium (USENIX Security 21) . USENIX Association, Aug. 2021, pp. 1559–1575
2021
Cited alongside, same era.
C. Chen and J. Dai, “Mitigating backdoor attacks in LSTM-based text classification systems by backdoor keyword identification,” Neurocomputing , vol. 452, pp. 253–262, 2021
2021
Cited alongside, same era.
F. Qi, Y. Chen, M. Li, Y. Yao, Z. Liu, and M. Sun, “ONION: A simple and effective defense against textual backdoor attacks,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Online and Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 9558–9566
2021
Cited alongside, same era.
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang, G. Li, L. Zhou, L. Shou, L. Zhou, M. Tufano, M. GONG, M. Zhou, N. Duan, N. Sundaresan, S. K. Deng, S. Fu, and S. LIU, “CodeXGLUE: A machine learning benchmark dataset for code understanding and generation,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1) , 2021
2021
Cited alongside, same era.
Y. Wang, W. Wang, S. Joty, and S. C. Hoi, “CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Online and Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 8696–8708
2021
Cited alongside, same era.
W. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang, “Unified pre-training for program understanding and generation,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Online: Association for Computational Linguistics, Jun. 2021, pp. 2655–2668
2021
Cited alongside, same era.
X. Chen, A. Salem, D. Chen, M. Backes, S. Ma, Q. Shen, Z. Wu, and Y. Zhang, “BadNL: Backdoor attacks against nlp models with semantic-preserving improvements,” in Annual computer security applications conference , 2021, pp. 554–569
2021
Cited alongside, same era.
Y. Li, X. Lyu, N. Koren, L. Lyu, B. Li, and X. Ma, “Anti-backdoor learning: Training clean models on poisoned data,” Advances in Neural Information Processing Systems , vol. 34, pp. 14 900–14 912, 2021
2021
Cited alongside, same era.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Advances in neural information processing systems , vol. 35, pp. 22 199–22 213, 2022
2022
Cited alongside, same era.
2023
Closest in time.
Google, “DIDACT: Large sequence models for software development activities,” 2023, accessed on October 2023. [Online]. Available: https://blog.research.google/2023/05/large-sequence-models-for-software.html
2023
Closest in time.
GitHub, OpenAI, “GitHub Copilot - Your AI pair programmer,” 2021, accessed on October 2023. [Online]. Available: https://github.com/features/copilot/
2023
Closest in time.
Amazon, “AI Code Generator - Amazon CodeWhisperer,” 2023, accessed on October 2023. [Online]. Available: https://aws.amazon.com/codewhisperer/
2023
Closest in time.
A. Jha and C. K. Reddy, “Codeattack: Code-based adversarial attacks for pre-trained programming language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 12, 2023, pp. 14 892–14 900
2023
Closest in time.
Y. Li, S. Liu, K. Chen, X. Xie, T. Zhang, and Y. Liu, “Multi-target backdoor attacks for code pre-trained models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 7236–7254
2023
Closest in time.
M. R. I. Rabin, A. Hussain, M. A. Alipour, and V. J. Hellendoorn, “Memorization and generalization in neural code intelligence models,” Information and Software Technology (IST) , vol. 153(107066), pp. 1–20, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Yang, B. Xu, J. M. Zhang, H. J. Kang, J. Shi, J. He, and D. Lo, “Stealthy backdoor attack for code models,” 2023
2023
Closest in time.
2023
Closest in time.
W. Sun, Y. Chen, G. Tao, C. Fang, X. Zhang, Q. Zhang, and B. Luo, “Backdooring neural code search,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 9692–9708
2023
Closest in time.
2023
Closest in time.
Z. Liu, B. Shen, Z. Lin, F. Wang, and W. Wang, “Maximum entropy loss, the silver bullet targeting backdoor attacks in pre-trained language models,” in Findings of the Association for Computational Linguistics: ACL 2023 . Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 3850–3868
2023
Closest in time.
L. Pang, T. Sun, H. Ling, and C. Chen, “Removing backdoor behaviors with unlabeled data,” in ICLR 2023 Workshop on Backdoor Attacks and Defenses in Machine Learning , 2023
2023
Closest in time.