Evaluating large language models trained on code
Original
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. (2021) · 2021
Cited alongside, same era.
Solving linear algebra by program synthesis
Original
Drori, I. and Verma, N. (2021) · 2021
Cited alongside, same era.
Introducing Amazon CodeWhisperer, the ML-powered coding companion
Ankur, D. and Atul, D. (2022) · 2022
Cited alongside, same era.
Programming is hard–or at least it used to be: Educational opportunities and challenges of ai code generation
Original
Becker, B. A., Denny, P., Finnie-Ansley, J., Luxton-Reilly, A., Prather, J., and Santos, E. A. (2022) · 2022
Cited alongside, same era.
Fooling moss detection with pretrained language models
Biderman, S. R. and Raff, E. (2022) · 2022
Cited alongside, same era.
GPT takes the bar exam
Original
Bommarito II, M. and Katz, D. M. (2022) · 2022
Cited alongside, same era.
Conversing with Copilot: Exploring prompt engineering for solving cs1 problems using natural language
Original
Denny, P., Kumar, V., and Giacaman, N. (2022) · 2022
Cited alongside, same era.
The robots are coming: Exploring the implications of OpenAI Codex on introductory programming
Finnie-Ansley, J., Denny, P., Becker, B. A., Luxton-Reilly, A., and Prather, J. (2022) · 2022
Cited alongside, same era.
How well does chatgpt do when taking the medical licensing exams? the implications of large language models for medical education and knowledge assessment
Gilson, A., Safranek, C. W., Huang, T., Socrates, V., Chi, L. S., Taylor, R. A., and Chartash, D. (2022) · 2022
Cited alongside, same era.
Codex hacks HackerRank: Memorization issues and a framework for code synthesis evaluation
Original
Karmakar, A., Prenner, J. A., D’Ambros, M., and Robbes, R. (2022) · 2022
Cited alongside, same era.
Performance of ChatGPT on USMLE: Potential for ai-assisted medical education using large language models
Kung, T. H., Cheatham, M., Medinilla, A., Sillos, C., De Leon, L., Elepano, C., Madriaga, M., Aggabao, R., Diaz-Candido, G., Maningo, J., et al. (2022) · 2022
Cited alongside, same era.
Competition-level code generation with AlphaCode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Lago, A. D., Hubert, T., Choy, P., de Masson d’Autume, C., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D. J., Robson, E. S., Kohli, P., de Freitas, N., Kavukcuoglu, K., and Vinyals, O. (2022) · 2022
Cited alongside, same era.