Fetching the paper…
Reading the bibliography…
Automated code generation with large language models has gained significant traction, but there remains no guarantee on the correctness of generated code.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Dafny: An automatic program verifier for functional correctness
Leino, K. R. M · 2010
Earlier work this paper cites.
Proverbot 9000 : Neural networks for proof assistance, 2016
Redmon, J. and Sanchez-Stern, A · 2016
Earlier work this paper cites.
Dependent types and multi-monadic effects in F*
Swamy, N., Hriţcu, C., Keller, C., Rastogi, A., Delignat-Lavaud, A., Forest, S., Bhargavan, K., Fournet, C., Strub, P.-Y., Kohlweiss, M., Zinzindohoué, J.-K., and Zanella-Béguelin, S · 2016
Earlier work this paper cites.
Reinforcement learning of theorem proving
Kaliszyk, C., Urban, J., Michalewski, H., and Olšák, M · 2018
Earlier work this paper cites.
The Coq Proof Assistant
Coq Development Team · 2020
Earlier work this paper cites.
Tactok: semantics-aware proof synthesis
First, E., Brun, Y., and Guha, A · 2020
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Polu, S. and Sutskever, I · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Asleep at the keyboard? Assessing the security of GitHub Copilot’s code contributions
Pearce, H. A., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R · 2021
Earlier work this paper cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
Pan, A., Bhatia, K., and Steinhardt, J · 2022
Earlier work this paper cites.
Formal mathematics statement curriculum learning
Polu, S., Han, J. M., Zheng, K., Baksys, M., Babuschkin, I., and Sutskever, I · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning, 2022
Zelikman, E., Wu, Y., Mu, J., and Goodman, N. D · 2022
Earlier work this paper cites.
Let’s sample step by step: Adaptive-consistency for efficient reasoning and coding with llms, 2023
Aggarwal, P., Madaan, A., Yang, Y., and Mausam · 2023
Earlier work this paper cites.
Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs
Akyürek, A. F., Akyürek, E., Madaan, A., Kalyan, A., Clark, P., Wijaya, D., and Tandon, N · 2023
Earlier work this paper cites.
Teaching large language models to self-debug
Chen, X., Lin, M., Schärli, N., and Zhou, D · 2023
Earlier work this paper cites.
Chern, I., Chern, S., Chen, S., Yuan, W., Feng, K., Zhou, C., He, J., Neubig, G., Liu, P., et al · 2023
Earlier work this paper cites.
Baldur: Whole-proof generation and repair with large language models
First, E., Rabe, M., Ringer, T., and Brun, Y · 2023
Earlier work this paper cites.
Understanding the limits of AI coding
Hendler, J · 2023
Cited alongside, same era.
Large language models and simple, stupid bugs
Jesse, K., Ahmed, T., Devanbu, P., and Morgan, E · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world GitHub issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Cited alongside, same era.
Verus: Verifying Rust programs using linear ghost types
Lattuada, A., Hance, T., Cho, C., Brun, M., Subasinghe, I., Zhou, Y., Howell, J., Parno, B., and Hawblitzel, C · 2023
Cited alongside, same era.
Starcoder: may the source be with you!
Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., Liu, Q., Zheltonozhskii, E., Zhuo, T. Y., Wang, T., Dehaene, O., Davaadorj, M., Lamy-Poirier, J., Monteiro, J., Shliazhko, O., Gontier, N., Meade, N., Zebaze, A., Yee, M.-H., Umapathi, L. K., Zhu, J., Lipkin, B., Oblokulov, M., Wang, Z., Murthy, R., Stillerman, J., Patel, S. S., Abulkhanov, D., Zocca, M., Dey, M., Zhang, Z., Fahmy, N., Bhattacharyya, U., Yu, W., Singh, S., Luccioni, S., Villegas, P., Kunakov, M., Zhdanov, F., Romero, M., Lee, T., Timor, N., Ding, J., Schlesinger, C., Schoelkopf, H., Ebert, J., Dao, T., Mishra, M., Gu, A., Robinson, J., Anderson, C. J., Dolan-Gavitt, B., Contractor, D., Reddy, S., Fried, D., Bahdanau, D., Jernite, Y., Ferrandis, C. M., Hughes, S., Wolf, T., Guha, A., von Werra, L., and de Vries, H · 2023
Occasionally secure: A comparative analysis of code generation assistants
Elgedawy, R., Sadik, J., Dutta, S., Gautam, A., Georgiou, K., Gholamrezae, F., Ji, F., Lim, K., Liu, Q., and Ruoti, S · 2024
Closest in time.
Critic: Large language models can self-correct with tool-interactive critiquing, 2024
Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., and Chen, W · 2024
Closest in time.
V-star: Training verifiers for self-taught reasoners, 2024
Hosseini, A., Yuan, X., Malkin, N., Courville, A., Sordoni, A., and Agarwal, R · 2024
Closest in time.
When can LLMs actually correct their own mistakes? A critical survey of self-correction of LLMs
Kamoi, R., Zhang, Y., Zhang, N., Han, J., and Zhang, R · 2024
Closest in time.
A survey on deep learning for theorem proving
Li, Z., Sun, J., Murphy, L., Su, Q., Li, Z., Zhang, X., Yang, K., and Si, X · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Liu, J., Xia, C. S., Wang, Y., and Zhang, L · 2023
Cited alongside, same era.
A survey of deep learning for mathematical reasoning
Lu, P., Qiu, L., Yu, W., Welleck, S., and Chang, K.-W · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., and Clark, P · 2023
Cited alongside, same era.
Peng, B., Galley, M., He, P., Cheng, H., Xie, Y., Hu, Y., Huang, Q., Liden, L., Yu, Z., Chen, W., et al · 2023
Cited alongside, same era.
Do users write more insecure code with ai assistants?
Perry, N., Srivastava, M., Kumar, D., and Boneh, D · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools, 2023
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Cited alongside, same era.
Clover: Closed-loop verifiable code generation
Sun, C., Sheng, Y., Padon, O., and Barrett, C · 2023
Cited alongside, same era.
Closest in time.
Lean-star: Learning to interleave thinking and proving
Lin, H., Sun, Z., Yang, Y., and Welleck, S · 2024
Closest in time.
minicodeprops: a minimal benchmark for proving code properties
Lohn, E. and Welleck, S · 2024
Closest in time.
Dafnybench: A benchmark for formal software verification, 2024
Loughridge, C., Sun, Q., Ahrenbach, S., Cassano, F., Sun, C., Sheng, Y., Mudide, A., Misu, M. R. H., Amin, N., and Tegmark, M · 2024
Closest in time.
Towards ai-assisted synthesis of verified dafny methods
Misu, M. R. H., Lopes, C. V., Ma, I., and Noble, J · 2024
Closest in time.
Code llama: Open foundation models for code
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Closest in time.
Easy-to-hard generalization: Scalable alignment beyond human supervision
Sun, Z., Yu, L., Shen, Y., Liu, W., Yang, Y., Welleck, S., and Gan, C · 2024
Closest in time.
Qwen2.5: A party of foundation models
Team, Q · 2024
Closest in time.
Humaneval-verus: Hand-written examples of verified verus code derived from humaneval, 2024
The HumanEval-Verus Contributors · 2024
Closest in time.
Wang, T., Kulikov, I., Golovneva, O., Yu, P., Yuan, W., Dwivedi-Yu, J., Pang, R. Y., Fazel-Zarandi, M., Weston, J., and Li, X · 2024
Closest in time.
From decoding to meta-generation: Inference-time algorithms for large language models
Welleck, S., Bertsch, A., Finlayson, M., Schoelkopf, H., Xie, A., Neubig, G., Kulikov, I., and Harchaoui, Z · 2024
Closest in time.
AutoVerus: Automated proof generation for Rust code, 2024
Yang, C., Li, X., Misu, M. R. H., Yao, J., Cui, W., Gong, Y., Hawblitzel, C., Lahiri, S., Lorch, J. R., Lu, S., Yang, F., Zhou, Z., and Lu, S · 2024
Closest in time.
Tree of thoughts: deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2024
Closest in time.
Sglang: Efficient execution of structured language model programs, 2024
Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., Barrett, C., and Sheng, Y · 2024
Closest in time.