Fetching the paper…
Reading the bibliography…
Pre-trained Generative Language models (e.g.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 1909
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?. In Proceedings of the thirteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 201–208
Dumitru Erhan, Aaron Courville, Yoshua Bengio, and Pascal Vincent. 2010 · 2010
Earlier work this paper cites.
On the naturalness of software. In 2012 34th International Conference on Software Engineering (ICSE) . IEEE, 837–847
Abram Hindle, Earl T Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. 2012 · 2012
Earlier work this paper cites.
Code completion with statistical language models. In Acm Sigplan Notices , Vol. 49. ACM, 419–428
Veselin Raychev, Martin Vechev, and Eran Yahav. 2014 · 2014
Earlier work this paper cites.
Suggesting accurate method and class names. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering . 38–49
Miltiadis Allamanis, Earl T Barr, Christian Bird, and Charles Sutton. 2015 · 2015
Earlier work this paper cites.
An embarrassingly simple approach to zero-shot learning. In International conference on machine learning . PMLR, 2152–2161
Bernardino Romera-Paredes and Philip Torr. 2015 · 2015
Earlier work this paper cites.
On the naturalness of software
Abram Hindle, Earl T Barr, Mark Gabel, Zhendong Su, and Premkumar Devanbu. 2016 · 2016
Earlier work this paper cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle. 2016 · 2016
Earlier work this paper cites.
On the" naturalness" of buggy code. In 2016 IEEE/ACM 38th International Conference on Software Engineering (ICSE) . IEEE, 428–439
Baishakhi Ray, Vincent Hellendoorn, Saheel Godhane, Zhaopeng Tu, Alberto Bacchelli, and Premkumar Devanbu. 2016 · 2016
Earlier work this paper cites.
Deep learning code fragments for code clone detection. In Proceedings of the 31st IEEE/ACM International Conference on Automated Software Engineering . ACM, 87–98
Martin White, Michele Tufano, Christopher Vendome, and Denys Poshyvanyk. 2016 · 2016
Earlier work this paper cites.
Learning to represent programs with graphs
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi. 2017 · 2017
Earlier work this paper cites.
Neural attribute machines for program generation
Matthew Amodio, Swarat Chaudhuri, and Thomas W Reps. 2017 · 2017
Earlier work this paper cites.
DeepFix: Fixing Common C Language Errors by Deep Learning.. In AAAI . 1345–1351
Rahul Gupta, Soham Pal, Aditya Kanade, and Shirish Shevade. 2017 · 2017
Earlier work this paper cites.
Are deep neural networks the best choice for modeling source code?. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering . ACM, 763–773
Vincent J Hellendoorn and Premkumar Devanbu. 2017 · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Recovering clear, natural identifiers from obfuscated JS names. In Proceedings of the 2017 11th joint meeting on foundations of software engineering . 683–693
Bogdan Vasilescu, Casey Casalnuovo, and Premkumar Devanbu. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems 30 . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
A survey of machine learning for big code and naturalness
Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton. 2018 · 2018
Earlier work this paper cites.
Refactoring: improving the design of existing code
Martin Fowler. 2018 · 2018
Cited alongside, same era.
Prevalence of confusing code in software projects: Atoms of confusion in the wild. In Proceedings of the 15th International Conference on Mining Software Repositories . 281–291
Dan Gopstein, Hongwei Henry Zhou, Phyllis Frankl, and Justin Cappos. 2018 · 2018
Cited alongside, same era.
Mapping language to code in programmatic context
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly
Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata. 2018 · 2018
Cited alongside, same era.
The adverse effects of code duplication in machine learning models of code. In Proceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software . 143–153
Code to Comment “Translation”: Data, Metrics, Baselining & Evaluation. In 2020 35th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 746–757
David Gros, Hariharan Sezhiyan, Prem Devanbu, and Zhou Yu. 2020 · 2020
Later among the works it cites.
Unsupervised Translation of Programming Languages
Marie-Anne Lachaux, Baptiste Roziere, Lowik Chanussot, and Guillaume Lample. 2020 · 2020
Later among the works it cites.
Retrieval-augmented generation for code summarization via hybrid gnn
Shangqing Liu, Yu Chen, Xiaofei Xie, Jingkai Siow, and Yang Liu. 2020 · 2020
Later among the works it cites.
Generalizing from a few examples: A survey on few-shot learning
Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni. 2020 · 2020
Later among the works it cites.
Unified Pre-training for Program Understanding and Generation. In 2021 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Miltiadis Allamanis. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. [n.d.] · 2019
Cited alongside, same era.
Coupling Retrieval and Meta-Learning for Context-Dependent Semantic Parsing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, Italy, 855–866
Daya Guo, Duyu Tang, Nan Duan, Ming Zhou, and Jian Yin. 2019 · 2019
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 2019
Cited alongside, same era.
Meta-transfer learning for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 403–412
Qianru Sun, Yaoyao Liu, Tat-Seng Chua, and Bernt Schiele. 2019 · 2019
Cited alongside, same era.
On Learning Meaningful Code Changes via Neural Machine Translation
Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk. 2019a · 2019
Cited alongside, same era.
Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021a · 2021
Later among the works it cites.
On Multi-Modal Learning of Editing Source Code. In 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE) . 443–455
Saikat Chakraborty and Baishakhi Ray. 2021 · 2021
Later among the works it cites.
Evaluating Large Language Models Trained on Code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Later among the works it cites.
Contrastive Learning for Source Code with Structural and Functional Properties
Yangruibo Ding, Luca Buratti, Saurabh Pujar, Alessandro Morari, Baishakhi Ray, and Saikat Chakraborty. 2021 · 2021
Later among the works it cites.
GraphCodeBERT: Pre-training Code Representations with Data Flow. In International Conference on Learning Representations
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Jian Yin, Daxin Jiang, et al · 2021
Later among the works it cites.
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Later among the works it cites.
Studying the usage of text-to-text transfer transformer to support code-related tasks. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 336–347
Antonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader Palacio, Denys Poshyvanyk, Rocco Oliveto, and Gabriele Bavota. 2021 · 2021
Later among the works it cites.
Retrieval Augmented Code Generation and Summarization
Md Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Later among the works it cites.
Semantic bug seeding: a learning-based approach for creating realistic bugs. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 906–918
Jibesh Patra and Michael Pradel. 2021 · 2021
Later among the works it cites.
CoTexT: Multi-task Learning with Code-Text Transformer
Long Phan, Hieu Tran, Daniel Le, Hieu Nguyen, James Anibal, Alec Peltekian, and Yanfang Ye. 2021 · 2021
Later among the works it cites.
Neural software analysis
Michael Pradel and Satish Chandra. 2021 · 2021
Later among the works it cites.
DOBF: A deobfuscation pre-training objective for programming languages
Baptiste Roziere, Marie-Anne Lachaux, Marc Szafraniec, and Guillaume Lample. 2021 · 2021
Later among the works it cites.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Later among the works it cites.
SPT-Code: Sequence-to-Sequence Pre-Training for Learning the Representation of Source Code
Changan Niu, Chuanyi Li, Vincent Ng, Jidong Ge, Liguo Huang, and Bin Luo. 2022 · 2022
Closest in time.
Diet Code is Healthy: Simplifying Programs for Pre-Trained Models of Code. In Proceedings of the 2022 The ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Singapore, Singapore) (ESEC/FSE 2022)
Zhaowei Zhang, Hongyu Zhang, Beijun Shen, and Xiaodong Gu. 2022 · 2022
Closest in time.
Summarizing Source Code using a Neural Attention Model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2073–2083
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016 · 2083
Closest in time.