Fetching the paper…
Reading the bibliography…
Pretrained language models have shown strong effectiveness in code-related tasks, such as code retrieval, code generation, code summarization, and code completion tasks.
Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of ACL . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ORANGE: a Method for Evaluating Automatic Evaluation Metrics for Machine Translation. In Proceedings of COLING . 501–507
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Finding clones with dup: Analysis of an experiment
Brenda S Baker. 2007 · 2007
Earlier work this paper cites.
An empirical study of function clones in open source software. In Proceedings of WCRE . 81–90
Chanchal K Roy and James R Cordy. 2008 · 2008
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Example-centric programming: integrating web search into the development environment. In Proceedings of CHI . 513–522
Joel Brandt, Mira Dontcheva, Marcos Weskamp, and Scott R. Klemmer. 2010 · 2010
Earlier work this paper cites.
Long short-term memory
Long Short-Term Memory. 2010 · 2010
Earlier work this paper cites.
Mining Source Code Repositories at Massive Scale using Language Modeling. In Proceedings of MSR . 207–216
Miltiadis Allamanis and Charles Sutton. 2013 · 2013
Earlier work this paper cites.
What help do developers seek, when and how?. In Proceedings of WCRE . 142–151
Hongwei Li, Zhenchang Xing, Xin Peng, and Wenyun Zhao. 2013 · 2013
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation. In Proceedings of EMNLP . 1412–1421
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
How developers search for code: a case study. In Proceedings of FSE . 191–201
Caitlin Sadowski, Kathryn T Stolee, and Sebastian Elbaum. 2015 · 2015
Earlier work this paper cites.
Evaluating clone detection tools with bigclonebench. In Proceedings of ICSME . 131–140
Jeffrey Svajlenko and Chanchal K Roy. 2015 · 2015
Earlier work this paper cites.
Probabilistic Model for Code with Decision Trees
Veselin Raychev, Pavol Bielik, and Martin Vechev. 2016 · 2016
Earlier work this paper cites.
Attention is All you Need. In Proceedings of NeurIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Mapping Language to Code in Programmatic Context. In Proceedings of EMNLP . 1643–1652
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Representation Degeneration Problem in Training Natural Language Generation Models. In Proceedings of ICLR
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019 · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of EMNLP . 3982–3992
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
ERNIE: Enhanced Language Representation with Informative Entities. In Proceedings of ACL . 1441–1451
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 2019
Earlier work this paper cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. In Proceedings of ICLR
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Earlier work this paper cites.
Cert: Contrastive self-supervised learning for language understanding
Hongchao Fang, Sicheng Wang, Meng Zhou, Jiayuan Ding, and Pengtao Xie. 2020 · 2020
Earlier work this paper cites.
CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Proceedings of EMNLP Findings . 1536–1547
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020 · 2020
Earlier work this paper cites.
Retrieval Augmented Language Model Pre-Training. In Proceedings of ICML . 3929–3938
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Cited alongside, same era.
CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2020 · 2020
Cited alongside, same era.
Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of EMNLP . 6769–6781
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of NeurIPS
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
On the Sentence Embeddings from Pre-trained Language Models. In Proceedings of EMNLP . 9119–9130
Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020 · 2020
Simple Entity-Centric Questions Challenge Dense Retrievers. In Proceedings of EMNLP . 6138–6148
Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021 · 2021
Later among the works it cites.
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of EMNLP . 8696–8708
Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi. 2021 · 2021
Later among the works it cites.
ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer. In Proceedings of ACL . 5065–5075
Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, and Weiran Xu. 2021 · 2021
Later among the works it cites.
Few-Shot Conversational Dense Retrieval. In Proceedings of SIGIR
Shi Yu, Zhenghao Liu, Chenyan Xiong, Tao Feng, and Zhiyuan Liu. 2021 · 2021
Later among the works it cites.
UniXcoder: Unified Cross-Modal Pre-training for Code Representation. In Proceedings of ACL . 7212–7225
Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2020
Cited alongside, same era.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Cited alongside, same era.
Clear: Contrastive learning for sentence representation
Zhuofeng Wu, Sinong Wang, Jiatao Gu, Madian Khabsa, Fei Sun, and Hao Ma. 2020 · 2020
Cited alongside, same era.
Coreferential Reasoning Learning for Language Representation. In Proceedings of EMNLP . 7170–7186
Deming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu, Peng Li, Maosong Sun, and Zhiyuan Liu. 2020 · 2020
Cited alongside, same era.
Unified Pre-training for Program Understanding and Generation. In Proceedings of NAACL-HLT . 2655–2668
Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Cited alongside, same era.
Program Synthesis with Large Language Models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code Search. In Proceedings of EMNLP . 2898–2910
Xiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, and Nan Duan. 2022 · 2022
Later among the works it cites.
ReACC: A Retrieval-Augmented Code Completion Framework. In Proceedings of ACL . 6227–6240
Shuai Lu, Nan Duan, Hojae Han, Daya Guo, Seung-won Hwang, and Alexey Svyatkovskiy. 2022 · 2022
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue
OpenAI. 2022 · 2022
Later among the works it cites.
CERT: Continual Pre-Training on Sketches for Library-Oriented Code Generation. In Proceedings of IJCAI
Daoguang Zan, Bei Chen, Dejian Yang, Zeqi Lin, Minsu Kim, Bei Guan, Yongji Wang, Weizhu Chen, and Jian-Guang Lou. 2022 · 2022
Later among the works it cites.
Docprompting: Generating code by retrieving the docs. In Proceedings of ICLR
Shuyan Zhou, Uri Alon, Frank F Xu, Zhengbao Jiang, and Graham Neubig. 2022 · 2022
Later among the works it cites.
Retrieval-Augmented Code Generation for Universal Information Extraction
Yucan Guo, Zixuan Li, Xiaolong Jin, Yantao Liu, Yutao Zeng, Wenxuan Liu, Xiang Li, Pan Yang, Long Bai, Jiafeng Guo, and Xueqi Cheng. 2023 · 2023
Later among the works it cites.
Active retrieval augmented generation
Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Context-Aware Code Generation Framework for Code Repositories: Local, Global, and Third-Party Library Awareness
Dianshu Liao, Shidong Pan, Qing Huang, Xiaoxue Ren, Zhenchang Xing, Huan Jin, and Qinying Li. 2023 · 2023
Later among the works it cites.
Universal Vision-Language Dense Retrieval: Learning A Unified Representation Space for Multi-Modal Retrieval. In Proceedings of ICLR
Zhenghao Liu, Chenyan Xiong, Yuanhuiyi Lv, Zhiyuan Liu, and Ge Yu. 2023 · 2023
Later among the works it cites.
SAIL: Search-Augmented Instruction Learning
Hongyin Luo, Yung-Sung Chuang, Yuan Gong, Tianhua Zhang, Yoon Kim, Xixin Wu, Danny Fox, Helen Meng, and James Glass. 2023 · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Later among the works it cites.
Entity-Augmented Code Generation
Anton Shapkin, Denis Litvinov, and Timofey Bryksin. 2023 · 2023
Later among the works it cites.
RepoFusion: Training Code Models to Understand Your Repository
Disha Shrivastava, Denis Kocetkov, Harm de Vries, Dzmitry Bahdanau, and Torsten Scholak. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
CodeT5+: Open Code Large Language Models for Code Understanding and Generation
Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi D. Q. Bui, Junnan Li, and Steven C. H. Hoi. 2023 · 2023
Later among the works it cites.
OpenMatch-v2: An All-in-One Multi-Modality PLM-Based Information Retrieval Toolkit. In Proceedings of SIGIR . 3160–3164
Shi Yu, Zhenghao Liu, Chenyan Xiong, and Zhiyuan Liu. 2023 · 2023
Later among the works it cites.
RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation
Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. 2023 · 2023
Later among the works it cites.
DeepSeek-Coder: When the Large Language Model Meets Programming – The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024 · 2024
Closest in time.
Code with CodeQwen1.5
Qwen Team. 2024 · 2024
Closest in time.