Fetching the paper…
Reading the bibliography…
Testing plays a crucial role in the software development cycle, enabling the detection of bugs, vulnerabilities, and other undesirable behaviors.
A complexity measure
Thomas J McCabe. 1976 · 1976
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown. 2020 · 2005
Earlier work this paper cites.
Unit test case generation with transformers and focal context
Michele Tufano, Dawn Drain, Alexey Svyatkovskiy, Shao Kun Deng, and Neel Sundaresan. 2020 · 2009
Earlier work this paper cites.
Testful: Automatic unit-test generation for java classes
Luciano Baresi and Matteo Miraz. 2010 · 2010
Earlier work this paper cites.
Mutation-driven generation of unit tests and oracles
Gordon Fraser and Andreas Zeller. 2010 · 2010
Earlier work this paper cites.
Symbolic execution for software testing in practice: preliminary assessment
Cristian Cadar, Patrice Godefroid, Sarfraz Khurshid, Corina S Păsăreanu, Koushik Sen, Nikolai Tillmann, and Willem Visser. 2011 · 2011
Earlier work this paper cites.
S2e: A platform for in-vivo multi-path analysis of software systems
Vitaly Chipounov, Volodymyr Kuznetsov, and George Candea. 2011 · 2011
Earlier work this paper cites.
Evosuite: automatic test suite generation for object-oriented software
Gordon Fraser and Andrea Arcuri. 2011 · 2011
Earlier work this paper cites.
A survey on unit testing practices and problems
Ermira Daka and Gordon Fraser. 2014 · 2014
Earlier work this paper cites.
Defects4j: A database of existing faults to enable controlled testing studies for java programs
René Just, Darioush Jalali, and Michael D Ernst. 2014 · 2014
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Earlier work this paper cites.
Automating code review activities by large-scale pre-training
Zhiyu Li, Shuai Lu, Daya Guo, Nan Duan, Shailesh Jannu, Grant Jenks, Deep Majumder, Jared Green, Alexey Svyatkovskiy, Shengyu Fu, et al. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
A3test: Assertion-augmented automated test case generation
Saranya Alagarsamy, Chakkrit Tantithamthavorn, and Aldeida Aleti. 2023 · 2023
Cited alongside, same era.
Longcoder: A long-range pre-trained language model for code completion
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley. 2023 · 2023
Cited alongside, same era.
Automated test case generation using code models and domain adaptation
Sepehr Hashtroudi, Jiho Shin, Hadi Hemmati, and Song Wang. 2023 · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2023 · 2023
Universal fuzzing via large language models
Chunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel, and Lingming Zhang. 2023 · 2023
Later among the works it cites.
Chatunitest: a chatgpt-based automated unit test generation tool
Zhuokui Xie, Yinghao Chen, Chen Zhi, Shuiguang Deng, and Jianwei Yin. 2023 · 2023
Later among the works it cites.
No more manual tests? evaluating and improving chatgpt for unit test generation
Zhiqiang Yuan, Yiling Lou, Mingwei Liu, Shiji Ding, Kaixin Wang, Yixuan Chen, and Xin Peng. 2023 · 2023
Later among the works it cites.
RepoCoder: Repository-level code completion through iterative retrieval and generation
Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. 2023a · 2023
Later among the works it cites.
Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ds-1000: A natural and reliable benchmark for data science code generation
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-tau Yih, Daniel Fried, Sida Wang, and Tao Yu. 2023 · 2023
Cited alongside, same era.
Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models
Caroline Lemieux, Jeevana Priya Inala, Shuvendu K Lahiri, and Siddhartha Sen. 2023 · 2023
Cited alongside, same era.
Prompting code interpreter to write better unit tests on quixbugs functions
Vincent Li and Nick Doiron. 2023 · 2023
Cited alongside, same era.
Repobench: Benchmarking repository-level code auto-completion systems
Tianyang Liu, Canwen Xu, and Julian McAuley. 2023 · 2023
Cited alongside, same era.
Cat-lm training language models on aligned code and tests
Nikitha Rao, Kush Jain, Uri Alon, Claire Le Goues, and Vincent J Hellendoorn. 2023 · 2023
Cited alongside, same era.
An empirical evaluation of using large language models for automated unit test generation
Max Schäfer, Sarah Nadi, Aryaz Eghbali, and Frank Tip. 2023 · 2023
Cited alongside, same era.
Reinforcement learning from automatic feedback for high-quality unit test generation
Benjamin Steenhoek, Michele Tufano, Neel Sundaresan, and Alexey Svyatkovskiy. 2023 · 2023
Cited alongside, same era.
Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, et al. 2023 · 2023
Later among the works it cites.
Effective test generation using pre-trained large language models and mutation testing
Arghavan Moradi Dakhel, Amin Nikanjam, Vahid Majdinasab, Foutse Khomh, and Michel C Desmarais. 2024 · 2024
Closest in time.
Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion
Yangruibo Ding, Zijian Wang, Wasi Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, et al. 2024 · 2024
Closest in time.
Cruxeval: A benchmark for code reasoning, understanding and execution
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I Wang. 2024 · 2024
Closest in time.
DevEval: A manually-annotated code generation benchmark aligned with real-world code repositories
Jia Li, Ge Li, Yunfei Zhao, Yongmin Li, Huanyu Liu, Hao Zhu, Lecheng Wang, Kaibo Liu, Zheng Fang, Lanshen Wang, Jiazheng Ding, Xuanming Zhang, Yuqi Zhu, Yihong Dong, Zhi Jin, Binhua Li, Fei Huang, Yongbin Li, Bin Gu, and Mengfei Yang. 2024b · 2024
Closest in time.
Coverup: Coverage-guided llm-based test generation
Juan Altmayer Pizzorno and Emery D Berger. 2024 · 2024
Closest in time.
Automatic generation of test cases based on bug reports: a feasibility study with large language models
Laura Plein, Wendkûuni C Ouédraogo, Jacques Klein, and Tegawendé F Bissyandé. 2024 · 2024
Closest in time.
Code-aware prompting: A study of coverage-guided test generation in regression setting using llm
Gabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang, Xiaofei Ma, Murali Krishna Ramanathan, and Baishakhi Ray. 2024 · 2024
Closest in time.