Fetching the paper…
Reading the bibliography…
Recently, fine-tuning pre-trained code models such as CodeBERT on downstream tasks has achieved great success in many software testing and analysis tasks.
CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 1909
Earlier work this paper cites.
A complexity measure
Thomas J McCabe. 1976 · 1976
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. 1994 · 1994
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation. In ACL . ACL, 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In ACL
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In IEEvaluation@ACL
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Compilers: principles, techniques, & tools
Alfred V Aho, Monica S Lam, Ravi Sethi, and Jeffrey D Ullman. 2007 · 2007
Earlier work this paper cites.
Representational similarity analysis-connecting the branches of systems neuroscience
Nikolaus Kriegeskorte, Marieke Mur, and Peter A Bandettini. 2008 · 2008
Earlier work this paper cites.
Christopher D. Manning, Prabhakar Raghavan, Hinrich Schütze, Introduction to Information Retrieval, Cambridge University Press. 2008. ISBN-13 978-0-521-86571-5, xxi+ 482 pages
Mark Sanderson. 2010 · 2010
Earlier work this paper cites.
Code clone detection experience at Microsoft. In Proceedings of the 5th International Workshop on Software Clones . 63–64
Yingnong Dang, Song Ge, Ray Huang, and Dongmei Zhang. 2011 · 2011
Earlier work this paper cites.
Static program analysis
Anders Møller and Michael I Schwartzbach. 2012 · 2012
Earlier work this paper cites.
Mining Source Code Repositories at Massive Scale using Language Modeling. In 2013 10th Working Conference on Mining Software Repositories (MSR) . IEEE, 207–216
Miltiadis Allamanis and Charles Sutton. 2013 · 2013
Earlier work this paper cites.
Towards a big data curated benchmark of inter-project code clones. In 2014 IEEE International Conference on Software Maintenance and Evolution . IEEE, 476–480
Jeffrey Svajlenko, Judith F Islam, Iman Keivanloo, Chanchal K Roy, and Mohammad Mamun Mia. 2014 · 2014
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation. In CVPR
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Convolutional neural networks over tree structures for programming language processing. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence . 1287–1293
Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin. 2016 · 2016
Earlier work this paper cites.
Attention is All you Need. In NIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
User’s guide to correlation coefficients
Haldun Akoglu. 2018 · 2018
Earlier work this paper cites.
Deep code search. In ICSE . ACM, 933–944
Xiaodong Gu, Hongyu Zhang, and Sunghun Kim. 2018 · 2018
Earlier work this paper cites.
Mapping language to code in programmatic context
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018a · 2018
Earlier work this paper cites.
Correlation coefficients: appropriate use and interpretation
Patrick Schober, Christa Boer, and Lothar A Schwarte. 2018 · 2018
Cited alongside, same era.
Blackbox Meets Blackbox: Representational Similarity & Stability Analysis of Neural Language Models and Brains. In BlackboxNLP@ACL . Association for Computational Linguistics, 191–203
Samira Abnar, Lisa Beinborn, Rochelle Choenni, and Willem H. Zuidema. 2019 · 2019
Cited alongside, same era.
Correlating Neural and Symbolic Representations of Language. In ACL (1) . Association for Computational Linguistics, 2952–2962
Grzegorz Chrupala and Afra Alishahi. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1) . Association for Computational Linguistics, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Parameter-Efficient Transfer Learning for NLP. In ICML (Proceedings of Machine Learning Research, Vol. 97) . PMLR, 2790–2799
CURE: Code-Aware Neural Machine Translation for Automatic Program Repair. In ICSE . IEEE, 1161–1173
Nan Jiang, Thibaud Lutellier, and Lin Tan. 2021 · 2021
Later among the works it cites.
What do pre-trained code models know about code?. In 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 1332–1336
Anjan Karmakar and Romain Robbes. 2021 · 2021
Later among the works it cites.
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021 · 2021
Later among the works it cites.
Automatic Program Repair with OpenAI’s Codex: Evaluating QuixBugs
Julian Aron Prenner and Romain Robbes. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In EMNLP (Findings)
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
A unified architecture for accelerating distributed { \{ DNN } \} training in heterogeneous { \{ GPU/CPU } \} clusters. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 463–479
Yimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi, Yong Cui, and Chuanxiong Guo. 2020 · 2020
Cited alongside, same era.
TinyBERT: Distilling BERT for Natural Language Understanding. In EMNLP (Findings) (Findings of ACL, Vol. EMNLP 2020) . Association for Computational Linguistics, 4163–4174
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Twinbert: Distilling knowledge to twin-structured compressed bert models for large-scale retrieval. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management . 2645–2652
Wenhao Lu, Jian Jiao, and Ruofei Zhang. 2020 · 2020
Cited alongside, same era.
What Happens To BERT Embeddings During Fine-tuning?. In BlackboxNLP@EMNLP . Association for Computational Linguistics, 33–44
Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020 · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In EMNLP (1) . Association for Computational Linguistics, 8696–8708
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021 · 2021
Later among the works it cites.
BEiT: BERT Pre-Training of Image Transformers. In ICLR . OpenReview.net
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. 2022 · 2022
Later among the works it cites.
NatGen: Generative pre-training by" Naturalizing" source code
Saikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar Devanbu, and Baishakhi Ray. 2022 · 2022
Later among the works it cites.
Towards Learning (Dis)-Similarity of Source Code from Program Contrasts. In ACL (1) . Association for Computational Linguistics, 6300–6312
Yangruibo Ding, Luca Buratti, Saurabh Pujar, Alessandro Morari, Baishakhi Ray, and Saikat Chakraborty. 2022 · 2022
Later among the works it cites.
UniXcoder: Unified Cross-Modal Pre-training for Code Representation. In ACL (1) . Association for Computational Linguistics, 7212–7225
Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022 · 2022
Later among the works it cites.
Masked Autoencoders Are Scalable Vision Learners. In CVPR . IEEE, 15979–15988
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross B. Girshick. 2022 · 2022
Later among the works it cites.
AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language models. In 37th IEEE/ACM International Conference on Automated Software Engineering . 1–11
José Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, and Houari Sahraoui. 2022 · 2022
Later among the works it cites.
Transferability in Deep Learning: A Survey
Junguang Jiang, Yang Shu, Jianmin Wang, and Mingsheng Long. 2022 · 2022
Later among the works it cites.
SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code Representations. In ICSE . ACM, 1–13
Changan Niu, Chuanyi Li, Vincent Ng, Jidong Ge, Liguo Huang, and Bin Luo. 2022 · 2022
Later among the works it cites.
Enhancing Semantic Code Search with Multimodal Contrastive Learning and Soft Data Augmentation
Ensheng Shi, Wenchao Gub, Yanlin Wang, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun. 2022a · 2022
Later among the works it cites.
Probing Pretrained Models of Source Code
Sergey Troshin and Nadezhda Chirkova. 2022 · 2022
Later among the works it cites.
Using Pre-Trained Models to Boost Code Review Automation. In ICSE . ACM, 2291–2302
Rosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella, Denys Poshyvanyk, and Gabriele Bavota. 2022 · 2022
Later among the works it cites.
What Do They Capture? - A Structural Analysis of Pre-Trained Language Models for Source Code. In ICSE . ACM, 2377–2388
Yao Wan, Wei Zhao, Hongyu Zhang, Yulei Sui, Guandong Xu, and Hai Jin. 2022 · 2022
Later among the works it cites.
Efficient DNN Training with Knowledge-Guided Layer Freezing
Yiding Wang, Decang Sun, Kai Chen, Fan Lai, and Mosharaf Chowdhury. 2022 · 2022
Later among the works it cites.
Zhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li, Yuqun Zhang, and Lingming Zhang. 2022 · 2022
Later among the works it cites.
Replication Package
Telly. 2023 · 2023
Closest in time.