Fetching the paper…
Reading the bibliography…
We present CRUXEval (Code Reasoning, Understanding, and eXecution Evaluation), a benchmark consisting of 800 Python functions (3-13 lines).
Summarizing source code using a neural attention model
Iyer, S., Konstas, I., Cheung, A., and Zettlemoyer, L · 2016
Earlier work this paper cites.
Barone, A. V. M. and Sennrich, R · 2017
Earlier work this paper cites.
Deepfix: Fixing common c language errors by deep learning
Gupta, R., Pal, S., Kanade, A., and Shevade, S · 2017
Earlier work this paper cites.
code2seq: Generating sequences from structured representations of code
Alon, U., Brody, S., Levy, O., and Yahav, E · 2018
Earlier work this paper cites.
Juice: A large scale distantly supervised dataset for open domain context-based code generation
Agashe, R., Iyer, S., and Zettlemoyer, L · 2019
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., and Brockschmidt, M · 2019
Earlier work this paper cites.
A neural model for generating natural language summaries of program subroutines
LeClair, A., Jiang, S., and McMillan, C · 2019
Earlier work this paper cites.
Nl2type: inferring javascript function types from natural language information
Malik, R. S., Patra, J., and Pradel, M · 2019
Earlier work this paper cites.
An empirical study on learning bug-fixing patches in the wild via neural machine translation
Tufano, M., Watson, C., Bavota, G., Penta, M. D., White, M., and Poshyvanyk, D · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Atom: Commit message generation based on abstract syntax tree and hybrid ranking
Liu, S., Gao, C., Chen, S., Nie, L. Y., and Liu, Y · 2020
Earlier work this paper cites.
Unsupervised translation of programming languages
Roziere, B., Lachaux, M.-A., Chanussot, L., and Lample, G · 2020
Earlier work this paper cites.
On learning meaningful assert statements for unit test cases
Watson, C., Tufano, M., Moran, K., Bavota, G., and Poshyvanyk, D · 2020
Earlier work this paper cites.
Avatar: A parallel corpus for java-python program translation
Ahmad, W. U., Tushar, M. G. R., Chakraborty, S., and Chang, K.-W · 2021
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Tfix: Learning to fix coding errors with a text-to-text transformer
Berabi, B., He, J., Raychev, V., and Vechev, M · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Codesc: A large code-description parallel dataset
Hasan, M., Muttaqueen, T., Ishtiaq, A. A., Mehrab, K. S., Haque, M. M. A., Hasan, T., Ahmad, W. U., Iqbal, A., and Shahriyar, R · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., et al · 2021
Earlier work this paper cites.
Understanding by understanding not: Modeling negation in language models
Hosseini, A., Reddy, S., Bahdanau, D., Hjelm, R. D., Sordoni, A., and Courville, A · 2021
Earlier work this paper cites.
Provable limitations of acquiring meaning from ungrounded form: What will future language models understand?
Merrill, W. C., Goldberg, Y., Schwartz, R., and Smith, N. A · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al · 2021
Earlier work this paper cites.
Multi-lingual evaluation of code generation models
Athiwaratkun, B., Gouda, S. K., Wang, Z., Li, X., Tian, Y., Tan, M., Ahmad, W. U., Wang, S., Sun, Q., Shang, M., et al · 2022
Earlier work this paper cites.
Multipl-e: A scalable and extensible approach to benchmarking neural code generation
Cassano, F., Gouwar, J., Nguyen, D., Nguyen, S., Phipps-Costin, L., Pinckney, D., Yee, M.-H., Zi, Y., Anderson, C. J., Feldman, M. Q., et al · 2022
Earlier work this paper cites.
Codet: Code generation with generated tests
Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J.-G., and Chen, W · 2022
Earlier work this paper cites.
Cocomic: Code completion by jointly modeling in-file and cross-file context
Ding, Y., Wang, Z., Ahmad, W. U., Ramanathan, M. K., Nallapati, R., Bhatia, P., Roth, D., and Xiang, B · 2022
Earlier work this paper cites.
Incoder: A generative model for code infilling and synthesis
Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., tau Yih, W., Zettlemoyer, L., and Lewis, M · 2022
Earlier work this paper cites.
Deepperf: A deep learning-based approach for improving software performance
Garg, S., Moghaddam, R. Z., Clement, C. B., Sundaresan, N., and Wu, C · 2022
Earlier work this paper cites.
Language models can teach themselves to program better
Haluptzok, P., Bowers, M., and Kalai, A. T · 2022
Earlier work this paper cites.
Fixeval: Execution-based evaluation of program fixes for competitive programming problems
Haque, M. M. A., Ahmad, W. U., Lourentzou, I., and Brown, C · 2022
Cited alongside, same era.
Jigsaw: Large language models meet program synthesis
Jain, N., Vaidyanath, S., Iyer, A., Natarajan, N., Parthasarathy, S., Rajamani, S., and Sharma, R · 2022
Cited alongside, same era.
I speak, you verify: Toward trustworthy neural program synthesis
Key, D., Li, W.-D., and Ellis, K · 2022
Cited alongside, same era.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Le, H., Wang, Y., Gotmare, A. D., Savarese, S., and Hoi, S. C. H · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al · 2022
Cited alongside, same era.
The false promise of imitating proprietary llms
Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P., Levine, S., and Song, D · 2023
Later among the works it cites.
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., et al · 2023
Later among the works it cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Later among the works it cites.
Inferfix: End-to-end program repair with llms
Jin, M., Shahriar, S., Tufano, M., Shi, X., Lu, S., Sundaresan, N., and Svyatkovskiy, A · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can we generate shellcodes via natural language? an empirical study
Liguori, P., Al-Hossami, E., Cotroneo, D., Natella, R., Cukic, B., and Shaikh, S · 2022
Cited alongside, same era.
Type4py: Practical deep similarity learning-based type inference for python
Mir, A. M., Latoškinas, E., Proksch, S., and Gousios, G · 2022
Cited alongside, same era.
Codegen: An open large language model for code with multi-turn program synthesis
Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R · 2022
Cited alongside, same era.
Natural language to code translation with execution
Shi, F., Fried, D., Ghazvininejad, M., Zettlemoyer, L., and Wang, S. I · 2022
Cited alongside, same era.
Methods2test: A dataset of focal methods mapped to test cases
Tufano, M., Deng, S. K., Sundaresan, N., and Svyatkovskiy, A · 2022
Cited alongside, same era.
Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., Yih, W.-t., Fried, D., Wang, S., and Yu, T · 2023
Later among the works it cites.
Teaching arithmetic to small transformers
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D · 2023
Later among the works it cites.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2023
Later among the works it cites.
The expresssive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2023
Later among the works it cites.
Miceli-Barone, A. V., Barez, F., Konstas, I., and Cohen, S. B · 2023
Later among the works it cites.
State of what art? a call for multi-prompt llm evaluation
Mizrahi, M., Kaplan, G., Malkin, D., Dror, R., Shahaf, D., and Stanovsky, G · 2023
Later among the works it cites.
Lever: Learning to verify language-to-code generation with execution
Ni, A., Iyer, S., Radev, D., Stoyanov, V., Yih, W.-t., Wang, S., and Lin, X. V · 2023
Later among the works it cites.
Gpt-4 technical report. arxiv 2303.08774
OpenAI, R · 2023
Later among the works it cites.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2023
Later among the works it cites.
Peng, B., Galley, M., He, P., Cheng, H., Xie, Y., Hu, Y., Huang, Q., Liden, L., Yu, Z., Chen, W., and Gao, J · 2023
Later among the works it cites.
Phind, 2023
Royzen, M., Wei, J., and Coleman, R · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al · 2023
Later among the works it cites.
Pangu-coder2: Boosting large language models for code with ranking feedback
Shen, B., Zhang, J., Chen, T., Zan, D., Geng, B., Fu, A., Zeng, M., Yu, A., Ji, J., Zhao, J., et al · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Schärli, N., and Zhou, D · 2023
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Shinn, N., Labash, B., and Gopinath, A · 2023
Later among the works it cites.
Repository-level prompt generation for large language models of code
Shrivastava, D., Larochelle, H., and Tarlow, D · 2023
Later among the works it cites.
Test-case-driven programming understanding in large language models for better code generation
Tian, Z. and Chen, J · 2023
Later among the works it cites.
Llmseceval: A dataset of natural language prompts for security evaluations
Tony, C., Mutas, M., Ferreyra, N. E. D., and Scandariato, R · 2023
Later among the works it cites.
Llms cannot find reasoning errors, but can correct them!
Tyen, G., Mansoor, H., Chen, P., Mak, T., and Cărbune, V · 2023
Later among the works it cites.
Typet5: Seq2seq type inference using static analysis
Wei, J., Durrett, G., and Dillig, I · 2023
Later among the works it cites.
Wu, Z., Qiu, L., Ross, A., Akyürek, E., Chen, B., Wang, B., Kim, N., Andreas, J., and Kim, Y · 2023
Later among the works it cites.
Large language models meet nl2code: A survey
Zan, D., Chen, B., Zhang, F., Lu, D., Wu, B., Guan, B., Yongji, W., and Lou, J.-G · 2023
Later among the works it cites.
Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x
Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Wang, Z., Shen, L., Wang, A., Li, Y., et al · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J., Bengio, S., and Nakkiran, P · 2023
Later among the works it cites.