Fetching the paper…
Reading the bibliography…
While there has been plenty of work on generating tests from existing code, there has been limited work on generating tests from issues.
Test driven development: By example
Beck, K · 2002
Earlier work this paper cites.
Defects4J: A database of existing faults to enable controlled testing studies for java programs
Just, R., Jalali, D., and Ernst, M. D · 2014
Earlier work this paper cites.
Codet: Code generation with generated tests
Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J.-G., and Chen, W · 2023
Earlier work this paper cites.
Large language models are few-shot testers: Exploring LLM-based general bug reproduction
Kang, S., Yoon, J., and Yoo, S · 2023
Earlier work this paper cites.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Earlier work this paper cites.
Instruction tuning for large language models: A survey
Zhang, S., Dong, L., Li, X., Zhang, S., Sun, X., Wang, S., Li, J., Hu, R., Zhang, T., Wu, F., et al · 2023
Earlier work this paper cites.
SWE-Bench+: Enhanced coding benchmark for LLMs
Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., and Wang, S · 2024
Cited alongside, same era.
CodeR: Issue resolving with multi-agent and task graphs
Chen, D., Lin, S., Zeng, M., Zan, D., Wang, J.-G., Cheshkov, A., Sun, J., Yu, H., Dong, G., Aliev, A., et al · 2024
Cited alongside, same era.
Introducing SWE-bench Verified, August 2024
Chowdhury, N., Aung, J., Shern, C. J., Jaffe, O., Sherburn, D., Starace, G., Mays, E., Dias, R., Aljubeh, M., Glaese, M., Jimenez, C. E., Yang, J., Liu, K., and Madry, A · 2024
Cited alongside, same era.
SWE-bench: Can language models resolve real-world GitHub issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Cited alongside, same era.
Llms as continuous learners: Improving the reproduction of defective code in software issues
Automatic generation of test cases based on bug reports: a feasibility study with large language models
Plein, L., Ouédraogo, W. C., Klein, J., and Bissyandé, T. F · 2024
Later among the works it cites.
SpecRover: Code intent extraction via LLMs
Ruan, H., Zhang, Y., and Roychoudhury, A · 2024
Later among the works it cites.
An empirical evaluation of using large language models for automated unit test generation
Schäfer, M., Nadi, S., Eghbali, A., and Tip, F · 2024
Later among the works it cites.
Agentless: Demystifying LLM-based software engineering agents
Xia, C. S., Deng, Y., Dunn, S., and Zhang, L · 2024
Later among the works it cites.
SWE-Agent: Agent-computer interfaces enable automated software engineering
Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., and Press, O · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, Y., Ma, Y., Cao, R., Li, B., Huang, F., Gu, X., and Li, Y · 2024
Cited alongside, same era.
SWT-bench: Testing and validating real-world bug-fixes with code agents
Mündler, N., Mueller, M. N., He, J., and Vechev, M · 2024
Cited alongside, same era.
AEGIS: An agent-based framework for general bug reproduction from issue descriptions
Wang, X., Gao, P., Meng, X., Peng, C., Hu, R., Lin, Y., and Gao, C
Cited in the paper.
OpenDevin: An open platform for AI software developers as generalist agents
Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., Pan, J., Song, Y., Li, B., Singh, J., Tran, H. H., Li, F., Ma, R., Zheng, M., Qian, B., Shao, Y., Muennighoff, N., Zhang, Y., Hui, B., Lin, J., Brennan, R., Peng, H., Ji, H., and Neubig, G
Cited in the paper.
Later among the works it cites.
CodeMonkeys: Scaling test-time compute for software engineering
Ehrlich, R., Brown, B., Juravsky, J., Clark, R., Ré, C., and Mirhoseini, A · 2025
Closest in time.