Fetching the paper…
Reading the bibliography…
We consider the well-known and important tasks of clone detection and information retrieval for source code.
Cross-lingual transfer learning for question answering
Chia-Hsuan Lee and Hung-Yi Lee. 2019 · 1907
Earlier work this paper cites.
A program for identifying duplicated code
Brenda S Baker. 1993 · 1993
Earlier work this paper cites.
Identifying similar code with program dependence graphs
J. Krinke. 2001 · 2001
Earlier work this paper cites.
Cross-lingual question answering using off-the-shelf machine translation
Kisuh Ahn, Beatrice Alex, Johan Bos, Tiphaine Dalmas, Jochen L Leidner, and Matthew B Smillie. 2004 · 2004
Earlier work this paper cites.
Cross-lingual question answering by answer translation
Johan Bos and Malvina Nissim. 2006 · 2006
Earlier work this paper cites.
Keyword translation accuracy and cross-lingual question answering inchinese and japanese
Teruko Mitamura, Mengqiu Wang, Hideki Shima, and Frank Lin. 2006 · 2006
Earlier work this paper cites.
Applying wikipedia’s multilingual knowledge to cross–lingual question answering
Sergio Ferrández, Antonio Toral, Oscar Ferrández, Antonio Ferrández, and Rafael Munoz. 2007 · 2007
Earlier work this paper cites.
A survey on software clone detection research
Chanchal Roy and James Cordy. 2007 · 2007
Earlier work this paper cites.
Linking e-mails and source code artifacts
Alberto Bacchelli, Michele Lanza, and Romain Robbes. 2010 · 2010
Earlier work this paper cites.
Example-centric programming: integrating web search into the development environment
Joel Brandt, Mira Dontcheva, Marcos Weskamp, and Scott R Klemmer. 2010 · 2010
Earlier work this paper cites.
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021 · 2010
Earlier work this paper cites.
Synthetic data augmentation for zero-shot cross-lingual question answering
Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot, Djamé Seddah, and Jacopo Staiano. 2020 · 2010
Earlier work this paper cites.
Multilingual synthetic question and answer generation for cross-lingual reading comprehension
Siamak Shakeri, Noah Constant, Mihir Sanjay Kale, and Linting Xue. 2020 · 2010
Earlier work this paper cites.
Searching connected api subgraph via text phrases
Wing-Kwan Chan, Hong Cheng, and David Lo. 2012 · 2012
Earlier work this paper cites.
Facilitating crowd sourced software engineering via stack overflow
Ohad Barzilay, Christoph Treude, and Alexey Zagalsky. 2013 · 2013
Earlier work this paper cites.
Mining stackoverflow to turn the ide into a self-confident programming prompter
Luca Ponzanelli, Gabriele Bavota, Massimiliano Di Penta, Rocco Oliveto, and Michele Lanza. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks over tree structures for programming language processing
Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Sourcerercc: Scaling code clone detection to big-code
Hitesh Sajnani, Vaibhav Saini, Jeffrey Svajlenko, Chanchal K Roy, and Cristina V Lopes. 2016 · 2016
Cited alongside, same era.
Nlp2code: Code snippet content assist via natural language tasks
Brock Angus Campbell and Christoph Treude. 2017 · 2017
Cited alongside, same era.
Cclearner: A deep learning-based clone detection approach
Liuqing Li, He Feng, Wenjie Zhuang, Na Meng, and Barbara Ryder. 2017 · 2017
Cited alongside, same era.
Towards semantic clone detection via probabilistic software modeling
Hannes Thaller, Lukas Linsbauer, and Alexander Egyed. 2020 · 2020
Later among the works it cites.
Detecting code clones with graph neural network and flow-augmented abstract syntax tree
Wenhan Wang, Ge Li, Bo Ma, Xin Xia, and Zhi Jin. 2020 · 2020
Later among the works it cites.
Unsupervised corpus aware language model pre-training for dense passage retrieval
Luyu Gao and Jamie Callan. 2021 · 2021
Later among the works it cites.
Cascaded fast and slow models for efficient semantic code search
Akhilesh Deepak Gotmare, Junnan Li, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Later among the works it cites.
GraphCodeBERT: Pre-training Code Representations with Data Flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hakam W Alomari and Matthew Stephan. 2018 · 2018
Cited alongside, same era.
Deep code search
Xiaodong Gu, Hongyu Zhang, and Sunghun Kim. 2018 · 2018
Cited alongside, same era.
Unsupervised cross-lingual transfer of word embedding spaces
Ruochen Xu, Yiming Yang, Naoki Otani, and Yuexin Wu. 2018 · 2018
Cited alongside, same era.
Cross-lingual natural language generation via pre-training
Zewen Chi, Li Dong, Furu Wei, Wenhui Wang, Xian-Ling Mao, and Heyan Huang. 2019 · 2019
Cited alongside, same era.
CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 2019
Cited alongside, same era.
Cross-lingual training for automatic question generation
Vishwajeet Kumar, Nitish Joshi, Arijit Mukherjee, Ganesh Ramakrishnan, and Preethi Jyothi. 2019 · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Codesc: A large code–description parallel dataset
Masum Hasan, Tanveer Muttaqueen, Abdullah Al Ishtiaq, Kazi Sajeed Mehrab, Md Mahim Anjum Haque, Tahmid Hasan, Wasi Ahmad, Anindya Iqbal, and Rifat Shahriyar. 2021 · 2021
Later among the works it cites.
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard and Édouard Grave. 2021 · 2021
Later among the works it cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
X-METRA-ADA: Cross-lingual meta-transfer learning adaptation to natural language understanding and question answering
Meryem M’hamdi, Doo Soon Kim, Franck Dernoncourt, Trung Bui, Xiang Ren, and Jonathan May. 2021 · 2021
Later among the works it cites.
SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation
Xin Wang, Yasheng Wang, Fei Mi, Pingyi Zhou, Yao Wan, Xiao Liu, Li Li, Hao Wu, Jin Liu, and Xin Jiang. 2021 · 2021
Later among the works it cites.
Improving zero-shot cross-lingual transfer for multilingual question answering over knowledge graph
Yucheng Zhou, Xiubo Geng, Tao Shen, Wenqiang Zhang, and Daxin Jiang. 2021 · 2021
Later among the works it cites.
Addressing leakage in self-supervised contextualized code retrieval
Johannes Villmow, Viola Campos, Adrian Ulges, and Ulrich Schwanecke. 2022 · 2022
Later among the works it cites.
Santacoder: don’t reach for the stars!
Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, Logesh Kumar Umapathi, Carolyn Jane Anderson, Yangtian Zi, Joel Lamy Poirier, Hailey Schoelkopf, Sergey Troshin, Dmitry Abulkhanov, Manuel Romero, Michael Lappert, Francesco De Toni, Bernardo García del Río, Qian Liu, Shamik Bose, Urvashi Bhattacharyya, Terry Yue Zhuo, Ian Yu, Paulo Villegas, Marco Zocca, Sourab Mangrulkar, David Lansky, Huu Nguyen, Danish Contractor, Luis Villa, Jia Li, Dzmitry Bahdanau, Yacine Jernite, Sean Hughes, Daniel Fried, Arjun Guha, Harm de Vries, and Leandro von Werra. 2023 · 2023
Closest in time.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023 · 2023
Closest in time.
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023 · 2023
Closest in time.
Deepseek-coder: When the large language model meets programming – the rise of code intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024 · 2024
Closest in time.