Fetching the paper…
Reading the bibliography…
We introduce SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at \$1 million USD total in real-world payouts.
End-to-end (e2e) testing and evaluation of high-assurance systems
Paul, R., Tsai, W. T., Chen, Y., Fan, C., Cao, Z., and Huang, H · 2006
Earlier work this paper cites.
Differentiating integration testing and unit testing
Brar, H. K. and Kaur, P. J · 2015
Earlier work this paper cites.
Freelancers in the software development process: A systematic mapping study
Gupta, V., Fernández-Crehuet, J. M., and Hanne, T · 2020
Earlier work this paper cites.
Program synthesis with large language models, 2021
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., and Sutton, C · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps, 2021
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation, 2021
Lu, S., Guo, D., Ren, S., Huang, J., Svyatkovskiy, A., Blanco, A., Clement, C., Drain, D., Jiang, D., Tang, D., Li, G., Zhou, L., Shou, L., Zhou, L., Tufano, M., Gong, M., Zhou, M., Duan, N., Sundaresan, N., Deng, S. K., Fu, S., and Liu, S · 2021
Earlier work this paper cites.
Ds-1000: A natural and reliable benchmark for data science code generation, 2022
Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., tau Yih, S. W., Fried, D., Wang, S., and Yu, T · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., Hubert, T., Choy, P., de Masson d’Autume, C., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D. J., Sutherland Robson, E., Kohli, P., de Freitas, N., Kavukcuoglu, K., and Vinyals, O · 2022
Earlier work this paper cites.
Issue #14958: zip/postcode validation error message not displayed for entering ‘,‘ on the home address screen
Expensify · 2023
Earlier work this paper cites.
Repobench: Benchmarking repository-level code auto-completion systems, 2023
Liu, T., Xu, C., and McAuley, J · 2023
Cited alongside, same era.
Codegen: An open large language model for code with multi-turn program synthesis, 2023
Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C · 2023
Cited alongside, same era.
Openai preparedness framework, 2023
OpenAI · 2023
Cited alongside, same era.
Goal driven discovery of distributional differences via language descriptions, 2023
Zhong, R., Zhang, P., Li, S., Ahn, J., Klein, D., and Steinhardt, J · 2023
Cited alongside, same era.
Responsible scaling policy
Anthropic · 2024
Cited alongside, same era.
Octopack: Instruction tuning code large language models, 2024
Muennighoff, N., Liu, Q., Zebaze, A., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., von Werra, L., and Longpre, S · 2024
Later among the works it cites.
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers, 2024
Si, C., Yang, D., and Hashimoto, T · 2024
Later among the works it cites.
Scicode: A research coding benchmark curated by scientists, 2024
Tian, M., Gao, L., Zhang, S. D., Chen, X., Fan, C., Guo, X., Haas, R., Ji, P., Krongchon, K., Li, Y., Liu, S., Luo, D., Ma, Y., Tong, H., Trinh, K., Tian, C., Wang, Z., Wu, B., Xiong, Y., Yin, S., Zhu, M., Lieret, K., Lu, Y., Liu, G., Du, Y., Tao, T., Press, O., Callan, J., Huerta, E., and Peng, H · 2024
Later among the works it cites.
Swe-bench multimodal: Do ai systems generalize to visual software domains?, 2024
Yang, J., Jimenez, C. E., Zhang, A. L., Lieret, K., Yang, J., Wu, X., Press, O., Muennighoff, N., Synnaeve, G., Narasimhan, K. R., Yang, D., Wang, S. I., and Press, O · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chowdhury, N., Aung, J., Shern, C. J., Jaffe, O., Sherburn, D., Starace, G., Mays, E., Dias, R., Aljubeh, M., Glaese, M., Jimenez, C. E., Yang, J., Liu, K., and Madry, A · 2024
Cited alongside, same era.
Introducing the frontier safety framework, 2024
DeepMind, G · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024
Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I · 2024
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?, 2024
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2024
Cited alongside, same era.
Concept induction: Analyzing unstructured text with high-level concepts using lloom
Lam, M. S., Teoh, J., Landay, J. A., Heer, J., and Bernstein, M. S · 2024
Cited alongside, same era.
Issue #25889: Dev: Share code avatar differs from profile avatar
Expensify
Cited in the paper.
Issue #41239: Add support for copy/pasting images on ios
Expensify
Cited in the paper.
Later among the works it cites.
Naturalcodebench: Examining coding performance mismatch on humaneval and natural user prompts, 2024
Zhang, S., Zhao, H., Liu, X., Zheng, Q., Qi, Z., Gu, X., Zhang, X., Dong, Y., and Tang, J · 2024
Later among the works it cites.
Zhuo, T. Y., Vu, M. C., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I. N. B., Zhan, H., He, J., Paul, I., Brunner, S., Gong, C., Hoang, T., Zebaze, A. R., Hong, X., Li, W.-D., Kaddour, J., Xu, M., Zhang, Z., Yadav, P., Jain, N., Gu, A., Cheng, Z., Liu, J., Liu, Q., Wang, Z., Lo, D., Hui, B., Muennighoff, N., Fried, D., Du, X., de Vries, H., and Werra, L. V · 2024
Later among the works it cites.
Playwright: Fast and reliable end-to-end testing for modern web apps
Microsoft · 2025
Closest in time.
OpenAI Authentication API Documentation , 2025
OpenAI · 2025
Closest in time.
Codeelo: Benchmarking competition-level code generation of llms with human-comparable elo ratings
Quan, S., Yang, J., Yu, B., Zheng, B., Liu, D., Yang, A., Ren, X., Gao, B., Miao, Y., Feng, Y., Wang, Z., Yang, J., Cui, Z., Fan, Y., Zhang, Y., Hui, B., Lin, J., and Qwen Team, A. G · 2025
Closest in time.