Fetching the paper…
Reading the bibliography…
A recent study by Ahmed and Devanbu reported that using a corpus of code written in multilingual datasets to fine-tune multilingual Pre-trained Language Models (PLMs) achieves higher performance as opposed to using a corpus of code written in just one programming language.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
The TREC-8 question answering track report. In
Ellen M Voorhees et al · 1999
Earlier work this paper cites.
CCFinder: A multilinguistic token-based code clone detection system for large scale source code
Toshihiro Kamiya, Shinji Kusumoto, and Katsuro Inoue. 2002 · 2002
Earlier work this paper cites.
Evaluating Web-based Question Answering Systems.. In
Dragomir R Radev, Hong Qi, Harris Wu, and Weiguo Fan. 2002 · 2002
Earlier work this paper cites.
Cclearner: A deep learning-based clone detection approach. In
Liuqing Li, He Feng, Wenjie Zhuang, Na Meng, and Barbara Ryder. 2017 · 2017
Earlier work this paper cites.
Newsroom: A Dataset of 1.3 Million Summaries with Diverse Extractive Strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Earlier work this paper cites.
Deepbugs: A learning approach to name-based bug detection
Michael Pradel and Koushik Sen. 2018 · 2018
Earlier work this paper cites.
Improving Automatic Source Code Summarization via Deep Reinforcement Learning. In
Yao Wan, Zhou Zhao, Min Yang, Guandong Xu, Haochao Ying, Jian Wu, and Philip S. Yu. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Better Word Embeddings by Disentangling Contextual n-Gram Information. In
Prakhar Gupta, Matteo Pagliardini, and Martin Jaggi. 2019 · 2019
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 2019
Earlier work this paper cites.
Pre-trained contextual embedding of source code
Aditya Kanade, Petros Maniatis, Gogul Balakrishnan, and Kensen Shi. 2019 · 2019
Earlier work this paper cites.
Text summarization with pretrained encoders. In
Yang Liu and Mirella Lapata. 2019 · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators. In
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2020
Earlier work this paper cites.
PyMT5: multi-mode translation of natural language and Python code with transformers
Colin B Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan. 2020 · 2020
Cited alongside, same era.
Codebert: A pre-trained model for programming and natural languages. In
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Cited alongside, same era.
Code to Comment “Translation”: Data, Metrics, Baselining & Evaluation. In
David Gros, Hariharan Sezhiyan, Prem Devanbu, and Zhou Yu. 2020 · 2020
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding. In
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Learning and evaluating contextual embedding of source code. In
Aditya Kanade, Petros Maniatis, Gogul Balakrishnan, and Kensen Shi. 2020 · 2020
Cited alongside, same era.
Stack Overflow Developer Survey 2021
2021 · 2021
Later among the works it cites.
Unified Pre-training for Program Understanding and Generation. In
Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Later among the works it cites.
Multilingual training for Software Engineering
Toufique Ahmed and Premkumar Devanbu. 2021 · 2021
Later among the works it cites.
Program Synthesis with Large Language Models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Later among the works it cites.
Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving Transformations. In
Nghi DQ Bui, Yijun Yu, and Lingxiao Jiang. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rafael-Michael Karampatsis and Charles Sutton. 2020 · 2020
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Multi-task learning based pre-trained language model for code completion. In
Fang Liu, Ge Li, Yunfei Zhao, and Zhi Jin. 2020 · 2020
Cited alongside, same era.
Pre-trained models for natural language processing: A survey
Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Unsupervised Translation of Programming Languages
Baptiste Roziere, Marie-Anne Lachaux, Lowik Chanussot, and Guillaume Lample. 2020 · 2020
Cited alongside, same era.
Intellicode compose: Code generation using transformer. In
Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020 · 2020
Cited alongside, same era.
An Empirical Study on the Usage of Transformer Models for Code Completion
Matteo Ciniselli, Nathan Cooper, Luca Pascarella, Antonio Mastropaolo, Emad Aghajani, Denys Poshyvanyk, Massimiliano Di Penta, and Gabriele Bavota. 2021 · 2021
Later among the works it cites.
Graphcodebert: Pre-training code representations with data flow. In
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2021
Later among the works it cites.
TreeBERT: A Tree-Based Pre-Trained Model for Programming Language. In
Xue Jiang, Zhuoran Zheng, Chen Lyu, Liang Li, and Lei Lyu. 2021 · 2021
Later among the works it cites.
EDITSUM: A Retrieve-and-Edit Framework for Source Code Summarization. In
Jia Li, Yongmin Li, Ge Li, Xing Hu, Xin Xia, and Zhi Jin. 2021 · 2021
Later among the works it cites.
Retrieval-Augmented Generation for Code Summarization via Hybrid GNN. In
Shangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow, and Yang Liu. 2021 · 2021
Later among the works it cites.
Code to Comment Translation: A Comparative Study on Model Effectiveness & Errors
Junayed Mahmud, Fahim Faisal, Raihan Islam Arnob, Antonios Anastasopoulos, and Kevin Moran. 2021 · 2021
Later among the works it cites.
Studying the usage of text-to-text transfer transformer to support code-related tasks. In
Antonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader Palacio, Denys Poshyvanyk, Rocco Oliveto, and Gabriele Bavota. 2021 · 2021
Later among the works it cites.
Evaluation Methodologies for Code Learning Tasks
Pengyu Nie, Jiyang Zhang, Junyi Jessy Li, Raymond J Mooney, and Milos Gligoric. 2021 · 2021
Later among the works it cites.
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Later among the works it cites.
SYNCOBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation
Xin Wang, Fei Mi Yasheng Wang, Pingyi Zhou, Yao Wan, Xiao Liu, Li Li, Hao Wu, Jin Liu, and Xin Jiang. 2022 · 2022
Closest in time.