Fetching the paper…
Reading the bibliography…
Humans write code in a fundamentally interactive manner and rely on constant execution feedback to correct errors, resolve ambiguities, and decompose tasks.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Defcon capture the flag: defending vulnerable code from intense attack
C. Cowan, S. Arnold, S. Beattie, C. Wright, and J. Viega · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
picoCTF, 2013
C. M. University · 2013
Earlier work this paper cites.
Docker: lightweight linux containers for consistent development and deployment
D. Merkel · 2014
Earlier work this paper cites.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Language to logical form with neural attention, 2016
L. Dong and M. Lapata · 2016
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning, 2017
V. Zhong, C. Xiong, and R. Socher · 2017
Earlier work this paper cites.
Leveraging grammar and reinforcement learning for neural program synthesis, 2018
R. Bunel, M. Hausknecht, J. Devlin, R. Singh, and P. Kohli · 2018
Earlier work this paper cites.
Execution-guided neural program synthesis
X. Chen, C. Liu, and D. X. Song · 2018
Earlier work this paper cites.
NL2Bash: A corpus and semantic parser for natural language interface to the linux operating system
X. V. Lin, C. Wang, L. Zettlemoyer, and M. D. Ernst · 2018
Earlier work this paper cites.
sqlite3mysql, 2018
K. Tusar · 2018
Earlier work this paper cites.
Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task
T. Yu, R. Zhang, K. Yang, M. Yasunaga, D. Wang, Z. Li, J. Ma, I. Li, Q. Yao, S. Roman, Z. Zhang, and D. Radev · 2018
Earlier work this paper cites.
Juice: A large scale distantly supervised dataset for open domain context-based code generation
R. Agashe, S. Iyer, and L. Zettlemoyer · 2019
Earlier work this paper cites.
Write, execute, assess: Program synthesis with a repl, 2019
K. Ellis, M. Nye, Y. Pu, F. Sosa, J. Tenenbaum, and A. Solar-Lezama · 2019
Earlier work this paper cites.
Reranking for neural semantic parsing
P. Yin and G. Neubig · 2019
Earlier work this paper cites.
Pymt5: multi-mode translation of natural language and python code with transformers, 2020
C. B. Clement, D. Drain, J. Timcheck, A. Svyatkovskiy, and N. Sundaresan · 2020
Cited alongside, same era.
Codebert: A pre-trained model for programming and natural languages, 2020
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling, 2020
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy · 2020
Cited alongside, same era.
Codesearchnet challenge: Evaluating the state of semantic code search, 2020
H. Husain, H.-H. Wu, T. Gazit, M. Allamanis, and M. Brockschmidt · 2020
Cited alongside, same era.
Keep calm and explore: Language models for action generation in text-based games
S. Yao, R. Rao, M. Hausknecht, and K. Narasimhan · 2020
Cited alongside, same era.
Ds-1000: A natural and reliable benchmark for data science code generation
Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, S. W. tau Yih, D. Fried, S. Wang, and T. Yu · 2022
Later among the works it cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
H. Le, Y. Wang, A. D. Gotmare, S. Savarese, and S. C. H. Hoi · 2022
Later among the works it cites.
Competition-level code generation with AlphaCode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. D. Lago, T. Hubert, P. Choy, C. de Masson d’Autume, I. Babuschkin, X. Chen, P.-S. Huang, J. Welbl, S. Gowal, A. Cherepanov, J. Molloy, D. J. Mankowitz, E. S. Robson, P. Kohli, N. de Freitas, K. Kavukcuoglu, and O. Vinyals · 2022
Later among the works it cites.
Natural language to code translation with execution, 2022
F. Shi, D. Fried, M. Ghazvininejad, L. Zettlemoyer, and S. I. Wang · 2022
Later among the works it cites.
Compilable neural code generation with compiler feedback, 2022
X. Wang, Y. Wang, Y. Wan, F. Mi, Y. Li, P. Zhou, J. Liu, H. Wu, X. Jiang, and Q. Liu · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neurips 2020 nlc2cmd competition: Translating natural language to bash commands
M. Agarwal, T. Chakraborti, Q. Fu, D. Gros, X. V. Lin, J. Maene, K. Talamadupula, Z. Teng, and J. White · 2021
Cited alongside, same era.
Program synthesis with large language models, 2021
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton · 2021
Cited alongside, same era.
Evaluating large language models trained on code, 2021
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Cited alongside, same era.
Measuring coding challenge competence with apps
D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
KaggleDBQA: Realistic evaluation of text-to-SQL parsers
C.-H. Lee, O. Polozov, and M. Richardson · 2021
Cited alongside, same era.
Codexglue: A machine learning benchmark dataset for code understanding and generation
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. B. Clement, D. Drain, D. Jiang, D. Tang, G. Li, L. Zhou, L. Shou, L. Zhou, M. Tufano, M. Gong, M. Zhou, N. Duan, N. Sundaresan, S. K. Deng, S. Fu, and S. Liu · 2021
Cited alongside, same era.
Codenet: A large-scale ai for code dataset for learning a diversity of coding tasks, 2021
R. Puri, D. S. Kung, G. Janssen, W. Zhang, G. Domeniconi, V. Zolotov, J. Dolby, J. Chen, M. Choudhury, L. Decker, V. Thost, L. Buratti, S. Pujar, S. Ramji, U. Finkler, S. Malaika, and F. Reiss · 2021
Cited alongside, same era.
Later among the works it cites.
Natural language to code generation in interactive data science notebooks, 2022
P. Yin, W.-D. Li, K. Xiao, A. Rao, Y. Wen, K. Shi, J. Howland, P. Bailey, M. Catasta, H. Michalewski, A. Polozov, and C. Sutton · 2022
Later among the works it cites.
N-best hypotheses reranking for text-to-sql systems, 2022
L. Zeng, S. H. K. Parthasarathi, and D. Hakkani-Tur · 2022
Later among the works it cites.
Coder reviewer reranking for code generation, 2022
T. Zhang, T. Yu, T. B. Hashimoto, M. Lewis, W. tau Yih, D. Fried, and S. I. Wang · 2022
Later among the works it cites.
Palm 2 technical report, 2023
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, and et al · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Closest in time.
Code as policies: Language model programs for embodied control, 2023
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Closest in time.
Lever: Learning to verify language-to-code generation with execution, 2023
A. Ni, S. Iyer, D. Radev, V. Stoyanov, W. tau Yih, S. I. Wang, and X. V. Lin · 2023
Closest in time.
Codegen: An open large language model for code with multi-turn program synthesis
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong · 2023
Closest in time.
Reflexion: Language agents with verbal reinforcement learning, 2023
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Closest in time.
Codebertscore: Evaluating code generation with pretrained models of code, 2023
S. Zhou, U. Alon, S. Agarwal, and G. Neubig · 2023
Closest in time.