Fetching the paper…
Reading the bibliography…
We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities.
Experiments with a heuristic compiler
Simon, H. A · 1963
Earlier work this paper cites.
Toward automatic program synthesis
Manna, Z. and Waldinger, R. J · 1971
Earlier work this paper cites.
Fault localization using execution slices and dataflow tests
Agrawal, H., Horgan, J. R., London, S., and Wong, W. E · 1995
Earlier work this paper cites.
Application of dynamic slicing in program debugging
Korel, B. and Rilling, J · 1997
Earlier work this paper cites.
Genetic programming III: Darwinian invention and problem solving , volume 3
Koza, J. R., Andre, D., Keane, M. A., and Bennett III, F. H · 1999
Earlier work this paper cites.
Lecture 3: Nondeterministic computation
Barrington, I. M. and Maciel, A · 2000
Earlier work this paper cites.
The economic impacts of inadequate infrastructure for software testing
Planning, S · 2002
Earlier work this paper cites.
Bugfix: A learning-based tool to assist developers in fixing bugs
Jeffrey, D., Feng, M., Gupta, N., and Gupta, R · 2009
Earlier work this paper cites.
Automating string processing in spreadsheets using input-output examples
Gulwani, S · 2011
Earlier work this paper cites.
The economics of software quality
Jones, C. and Bonsignour, O · 2011
Earlier work this paper cites.
A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each
Goues, C. L., Dewey-Vogt, M., Forrest, S., and Weimer, W · 2012
Earlier work this paper cites.
Spreadsheet data manipulation using examples
Gulwani, S., Harris, W. R., and Singh, R · 2012
Earlier work this paper cites.
On the naturalness of software
Hindle, A., Barr, E. T., Su, Z., Gabel, M., and Devanbu, P · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Temporal logics for hyperproperties
Clarkson, M. R., Finkbeiner, B., Koleini, M., Micinski, K. K., Rabe, M. N., and Sánchez, C · 2014
Earlier work this paper cites.
Generating sequences with recurrent neural networks, 2014
Graves, A · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Structured generative models of natural source code
Maddison, C. J. and Tarlow, D · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Zaremba, W. and Sutskever, I · 2014
Earlier work this paper cites.
Bimodal modelling of source code and natural language
Allamanis, M., Tarlow, D., Gordon, A., and Wei, Y · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V · 2015
Earlier work this paper cites.
General program synthesis benchmark suite
Helmuth, T. and Spector, L · 2015
Earlier work this paper cites.
Kaiser, Ł. and Sutskever, I · 2015
Earlier work this paper cites.
An analysis of patch plausibility and correctness for generate-and-validate patch generation systems
Qi, Z., Long, F., Achour, S., and Rinard, M · 2015
Earlier work this paper cites.
End-to-end memory networks, 2015
Sukhbaatar, S., Szlam, A., Weston, J., and Fergus, R · 2015
Earlier work this paper cites.
Memory networks, 2015
Weston, J., Chopra, S., and Bordes, A · 2015
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., Colmenarejo, S. G., Grefenstette, E., Ramalho, T., Agapiou, J., et al · 2016
Earlier work this paper cites.
Latent predictor networks for code generation
Ling, W., Blunsom, P., Grefenstette, E., Hermann, K. M., Kočiskỳ, T., Wang, F., and Senior, A · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Neural programmer-interpreters, 2016
Reed, S. and de Freitas, N · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
Van Oord, A., Kalchbrenner, N., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Deepcoder: Learning to write programs
Balog, M., Gaunt, A., Brockschmidt, M., Nowozin, S., and Tarlow, D · 2017
Earlier work this paper cites.
Barone, A. V. M. and Sennrich, R · 2017
Earlier work this paper cites.
The trouble with bias
Crawford, K · 2017
Earlier work this paper cites.
Visual dialog
Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Moura, J. M., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Robustfill: Neural program learning under noisy i/o
Devlin, J., Uesato, J., Bhupatiraju, S., Singh, R., rahman Mohamed, A., and Kohli, P · 2017
Earlier work this paper cites.
On the difficulty of benchmarking inductive program synthesis methods
Pantridge, E., Helmuth, T., McPhee, N. F., and Spector, L · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
A syntactic neural model for general-purpose code generation
Yin, P. and Neubig, G · 2017
Cited alongside, same era.
code2seq: Generating sequences from structured representations of code
Alon, U., Brody, S., Levy, O., and Yahav, E · 2018
Cited alongside, same era.
Clarifying ”ai alignment”
Christiano, P · 2018
Cited alongside, same era.
Protecting applications with automated software diversity, Sep 2018
Davis, B · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Open-sourcing gvisor, a sandboxed container runtime, 2018
Lacasse, N · 2018
Cited alongside, same era.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Unsupervised translation of programming languages
Lachaux, M.-A., Rozière, B., Chanussot, L., and Lample, G · 2020
Later among the works it cites.
What distinguishes great software engineers?
Li, P. L., Ko, A. J., and Begel, A · 2020
Later among the works it cites.
Recalibrating global data center energy-use estimates
Masanet, E., Shehabi, A., Lei, N., Smith, S., and Koomey, J · 2020
Later among the works it cites.
Backstabber’s knife collection: A review of open source software supply chain attacks, 2020
Ohm, M., Plate, H., Sykosch, A., and Meier, M · 2020
Later among the works it cites.
Python developers survey 2020 results, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Handbook of Applied Cryptography
Menezes, A., van Oorschot, P., and Vanstone, S · 2018
Cited alongside, same era.
Generating high fidelity images with subscale pixel networks and multidimensional upscaling, 2018
Menick, J. and Kalchbrenner, N · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Improving neural program synthesis with inferred execution traces
Shin, E. C., Polosukhin, I., and Song, D · 2018
Cited alongside, same era.
Python Software Foundation and JetBrains · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N. M., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Later among the works it cites.
Codebleu: a method for automatic evaluation of code synthesis
Ren, S., Guo, D., Lu, S., Zhou, L., Liu, S., Tang, D., Sundaresan, N., Zhou, M., Blanco, A., and Ma, S · 2020
Later among the works it cites.
Sourcefinder: Finding malware source-code from publicly available repositories in github
Rokon, M. O. F., Islam, R., Darki, A., Papalexakis, E. E., and Faloutsos, M · 2020
Later among the works it cites.
You autocomplete me: Poisoning vulnerabilities in neural code completion
Schuster, R., Song, C., Tromer, E., and Shmatikov, V · 2020
Later among the works it cites.
2020 developer survey, 2020
Stack Overflow · 2020
Later among the works it cites.
Learning to summarize from human feedback, 2020
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2020
Later among the works it cites.
Unit test case generation with transformers and focal context
Tufano, M., Drain, D., Svyatkovskiy, A., Deng, S. K., and Sundaresan, N · 2020
Later among the works it cites.
Persistent anti-muslim bias in large language models
Abid, A., Farooqi, M., and Zou, J · 2021
Closest in time.
Learning autocompletion from real-world datasets
Aye, G. A., Kim, S., and Li, H · 2021
Closest in time.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Closest in time.
GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow, 2021
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Closest in time.
Extracting training data from large language models
Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., and Raffel, C · 2021
Closest in time.
Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence
Crawford, K · 2021
Closest in time.
Generating bug-fixes using pretrained transformers
Drain, D., Wu, C., Svyatkovskiy, A., and Sundaresan, N · 2021
Closest in time.
Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses, 2021
Goldblum, M., Tsipras, D., Xie, C., Chen, X., Schwarzschild, A., Song, D., Madry, A., Li, B., and Goldstein, T · 2021
Closest in time.
Measuring coding challenge competence with apps
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., et al · 2021
Closest in time.
Kenton, Z., Everitt, T., Weidinger, L., Gabriel, I., Mikulik, V., and Irving, G · 2021
Closest in time.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Lu, S., Guo, D., Ren, S., Huang, J., Svyatkovskiy, A., Blanco, A., Clement, C., Drain, D., Jiang, D., Tang, D., Li, G., Zhou, L., Shou, L., Zhou, L., Tufano, M., Gong, M., Zhou, M., Duan, N., Sundaresan, N., Deng, S. K., Fu, S., and Liu, S · 2021
Closest in time.
15-1252.00 - software developers, 2021
O*NET · 2021
Closest in time.
Carbon emissions and large neural network training
Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., and Dean, J · 2021
Closest in time.
Learning compositional neural programs with recursive tree search and planning, 2021
Pierrot, T., Ligner, G., Reed, S., Sigaud, O., Perrin, N., Laterre, A., Kas, D., Beguir, K., and de Freitas, N · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Closest in time.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Closest in time.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., et al · 2021
Closest in time.
Women’s participation in open source software: A survey of the literature
Trinkenreich, B., Wiese, I., Sarma, A., Gerosa, M., and Steinmacher, I · 2021
Closest in time.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Closest in time.
Fun and dystopia with ai-based code generation using gpt-j-6b, June 2021
Woolf, M · 2021
Closest in time.
In-ide code generation from natural language: Promise and challenges
Xu, F. F., Vasilescu, B., and Neubig, G · 2021
Closest in time.
Merlot: Multimodal neural script knowledge models
Zellers, R., Lu, X., Hessel, J., Yu, Y., Park, J. S., Cao, J., Farhadi, A., and Choi, Y · 2021
Closest in time.
Calibrate before use: Improving few-shot performance of language models
Zhao, T. Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Closest in time.
A first look at rote learning in github copilot suggestions., Jun 2021
Ziegler, A · 2021
Closest in time.