Fetching the paper…
Reading the bibliography…
Modern language models (LMs) have gained widespread acceptance in everyday and professional contexts, particularly in programming.
Adam: a method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
A C/C++ code vulnerability dataset with code changes and CVE summaries
Fan, J., Li, Y., Wang, S., and Nguyen, T. N · 2020
Earlier work this paper cites.
SourceFinder: finding malware source-code from publicly available repositories in GitHub
Rokon, M. O. F., Islam, R., Darki, A., Papalexakis, E. E., and Faloutsos, M · 2020
Earlier work this paper cites.
Neural text generation with unlikelihood training
Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Earlier work this paper cites.
VUDENC: vulnerability detection with deep learning on a natural codebase for python
Wartschinski, L., Noller, Y., Vogel, T., Kehrer, T., and Grunske, L · 2021
Earlier work this paper cites.
Constitutional AI: harmlessness from AI feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Earlier work this paper cites.
LoRA: low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Competition-level code generation with AlphaCode
Li, Y., Choi, D. H., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Lago, A. D., et al · 2022
Earlier work this paper cites.
Truthfulqa: measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Asleep at the keyboard? assessing the security of GitHub Copilot’s code contributions
Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R · 2022
Cited alongside, same era.
SecurityEval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques
Siddiq, M. L. and Santos, J. C. S · 2022
Cited alongside, same era.
URL https://huggingface.co/datasets/codefuse-ai/Evol-instruction-66k
HuggingFace: codefuse-ai/Evol-instruction-66k, 2023 · 2023
Cited alongside, same era.
Product Anthropic, 2023
Anthropic · 2023
Cited alongside, same era.
StarCoder: may the source be with you!
Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., et al · 2023
Later among the works it cites.
WizardCoder: empowering code large language models with Evol-Instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2023
Later among the works it cites.
CWE: common weakness enumerations, 2023
MITRE · 2023
Later among the works it cites.
Octopack: Instruction tuning code large language models
Muennighoff, N., Liu, Q., Zebaze, A., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., von Werra, L., and Longpre, S · 2023
Later among the works it cites.
CodeGen: an open large language model for code with multi-turn program synthesis
Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Code alpaca: an instruction-following LLaMA model for code generation, 2023
Chaudhary, S · 2023
Cited alongside, same era.
Data quality for software vulnerability datasets
Croft, R., Babar, M. A., and Kholoosi, M. M · 2023
Cited alongside, same era.
difflib - Helpers for computing deltas, 2023
difflib · 2023
Cited alongside, same era.
We analyzed millions of ChatGPT user sessions: Visits are down 29% since may, programming assistance is 30% of use, 2023
Fishkin, R · 2023
Cited alongside, same era.
CodeQL - GitHub, 2023
GitHub · 2023
Cited alongside, same era.
Large language models for code: security hardening and adversarial testing
He, J. and Vechev, M · 2023
Cited alongside, same era.
Phi-2: the surprising power of small language models, 2023
Javaheripi, M. and Bubeck, S · 2023
Cited alongside, same era.
Later among the works it cites.
Introducing Gemini: our largest and most capable AI model, 2023
Pichai, S. and Hassabis, D · 2023
Later among the works it cites.
Code Llama: open foundation models for code
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al · 2023
Later among the works it cites.
Introducing Microsoft 365 Copilot - your copilot for work, 2023
Spataro, J · 2023
Later among the works it cites.
Llama 2: open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Self-Instruct: aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2023
Later among the works it cites.
CodeT5+: open code large language models for code understanding and generation
Wang, Y., Le, H., Gotmare, A., Bui, N. D. Q., Li, J., and Hoi, S. C. H · 2023
Later among the works it cites.
Magicoder: source code is all you need
Wei, Y., Wang, Z., Liu, J., Ding, Y., and Zhang, L · 2023
Later among the works it cites.
GitHub Copilot Chat now generally available for organizations and individuals, 2023
Zhao, S · 2023
Later among the works it cites.
LMSYS-Chat-1M: a large-scale real-world LLM conversation dataset
Zheng, L., Chiang, W., Sheng, Y., Li, T., Zhuang, S., Wu, Z., Zhuang, Y., Li, Z., Lin, Z., Xing, E. P., et al · 2023
Later among the works it cites.