Fetching the paper…
Reading the bibliography…
Code completion, a highly valuable topic in the software development domain, has been increasingly promoted for use by recent advances in large language models (LLMs).
Measuring nominal scale agreement among many raters
Joseph L Fleiss · 1971
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Metamorphic testing: a new approach for generating next test cases
Tsong Y Chen, Shing C Cheung, and Shiu Ming Yiu · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei jing Zhu · 2002
Earlier work this paper cites.
GPT2: Empirical slant delay model for radio space geodetic techniques
Klemens Lagler, Michael Schindelegger, Johannes Böhm, Hana Krásná, and Tobias Nilsson · 2013
Earlier work this paper cites.
Convolutional neural networks over tree structures for programming language processing
Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin · 2016
Earlier work this paper cites.
Fairness testing: testing software for discrimination
Sainyam Galhotra, Yuriy Brun, and Alexandra Meliou · 2017
Earlier work this paper cites.
DeepXplore: Automated whitebox testing of deep learning systems
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana · 2017
Earlier work this paper cites.
code2seq: Generating sequences from structured representations of code
Uri Alon, Shaked Brody, Omer Levy, and Eran Yahav · 2018
Earlier work this paper cites.
Identifying implementation bugs in machine learning based image classifiers using metamorphic testing
Anurag Dwarakanath, Manish Ahuja, Samarth Sikand, Raghotham M. Rao, R. P. Jagadeesh Chandra Bose, Neville Dubash, and Sanjay Podder · 2018
Earlier work this paper cites.
Malware classification with deep convolutional neural networks
Mahmoud Kalash, Mrigank Rochan, Noman Mohammed, Neil DB Bruce, Yang Wang, and Farkhund Iqbal · 2018
Earlier work this paper cites.
Vuldeepecker: A deep learning-based system for vulnerability detection
Zhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin, Sujuan Wang, Zhijun Deng, and Yuyi Zhong · 2018
Earlier work this paper cites.
Tensorfuzz: Debugging neural networks with coverage-guided fuzzing
Augustus Odena and Ian Goodfellow · 2018
Earlier work this paper cites.
DeepTest: Automated testing of deep-neural-network-driven autonomous cars
Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray · 2018
Earlier work this paper cites.
Automated directed fairness testing
Sakshi Udeshi, Pryanshu Arora, and Sudipta Chattopadhyay · 2018
Earlier work this paper cites.
DeepRoad: GAN-based Metamorphic Testing and Input Validation Framework for Autonomous Driving Systems
Mengshi Zhang, Yuqun Zhang, Lingming Zhang, Cong Liu, and Sarfraz Khurshid · 2018
Earlier work this paper cites.
code2vec: Learning distributed representations of code
Uri Alon, Meital Zilberstein, Omer Levy, and Eran Yahav · 2019
Earlier work this paper cites.
When deep learning met code search
José Cambronero, Hongyu Li, Seohyun Kim, Koushik Sen, and Satish Chandra · 2019
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Generating biased dataset for metamorphic testing of machine learning programs
Shin Nakajima and Tsong Yueh Chen · 2019
Earlier work this paper cites.
Adversarial sample detection for deep neural network through model mutation testing
Jingyi Wang, Guoliang Dong, Jun Sun, Xinyu Wang, and Peixin Zhang · 2019
Earlier work this paper cites.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu · 2019
Earlier work this paper cites.
Adversarial robustness for code
Pavol Bielik and Martin Vechev · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Earlier work this paper cites.
Machine translation testing via pathological invariance
Shashij Gupta, Pinjia He, Clara Meister, and Zhendong Su · 2020
Cited alongside, same era.
Structure-invariant testing for machine translation
Pinjia He, Clara Meister, and Zhendong Su · 2020
Cited alongside, same era.
Metamorphic testing and certified mitigation of fairness violations in nlp models
Pingchuan Ma, Shuai Wang, and Jin Liu · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of nlp models with checklist
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Cited alongside, same era.
Automatic testing and improvement of machine translation
Zeyu Sun, Jie M Zhang, Mark Harman, Mike Papadakis, and Lu Zhang · 2020
Cited alongside, same era.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2020
Probing model signal-awareness via prediction-preserving input minimization
Sahil Suneja, Yunhui Zheng, Yufan Zhuang, Jim A Laredo, and Alessandro Morari · 2021
Later among the works it cites.
To what extent do dnn-based image classification models make unreliable inferences?
Yongqiang Tian, Shiqing Ma, Ming Wen, Yepang Liu, Shing-Chi Cheung, and Xiangyu Zhang · 2021
Later among the works it cites.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi · 2021
Later among the works it cites.
Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi · 2021
Later among the works it cites.
Centris: A precise and scalable approach for identifying modified open-source software reuse
Seunghoon Woo, Sunghan Park, Seulbae Kim, Heejo Lee, and Hakjoo Oh · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Combining graph-based learning with automated data collection for code vulnerability detection
Huanting Wang, Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Songfang Huang, Dingyi Fang, Yansong Feng, Lizhong Bian, and Zheng Wang · 2020
Cited alongside, same era.
Metamorphic object insertion for testing object detection systems
Shuai Wang and Zhendong Su · 2020
Cited alongside, same era.
Adversarial examples for models of code
Noam Yefet, Uri Alon, and Eran Yahav · 2020
Cited alongside, same era.
Codecmr: Cross-modal retrieval for function-level binary source code matching
Zeping Yu, Wenxin Zheng, Jiaqi Wang, Qiyi Tang, Sen Nie, and Shi Wu · 2020
Cited alongside, same era.
Deepbillboard: Systematic physical-world testing of autonomous driving systems
Husheng Zhou, Wei Li, Yuankun Zhu, Yuqun Zhang, Bei Yu, Lingming Zhang, and Cong Liu · 2020
Cited alongside, same era.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Cited alongside, same era.
Later among the works it cites.
Perception matters: Detecting perception failures of vqa models using metamorphic testing
Yuanyuan Yuan, Shuai Wang, Mingyue Jiang, and Tsong Yueh Chen · 2021
Later among the works it cites.
Language-agnostic representation learning of source code from structure and context
Daniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec, and Stephan Günnemann · 2021
Later among the works it cites.
Fairness testing: A comprehensive survey and analysis of trends
Zhenpeng Chen, Jie M Zhang, Max Hort, Federica Sarro, and Mark Harman · 2022
Closest in time.
Unixcoder: Unified cross-modal pre-training for code representation
Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin · 2022
Closest in time.
Semantic robustness of models of source code
Jordan Henke, Goutham Ramakrishnan, Zi Wang, Aws Albarghouth, Somesh Jha, and Thomas Reps · 2022
Closest in time.
Is github copilot a substitute for human pair-programming? an empirical study
Saki Imai · 2022
Closest in time.
Deep learning application on code clone detection: A review of current knowledge
Maggie Lei, Hao Li, Ji Li, Namrata Aundhkar, and Dae-Kyoo Kim · 2022
Closest in time.
Competition-level code generation with alphacode
Yujia Li, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals · 2022
Closest in time.
Unleashing the power of compiler intermediate representation to enhance neural program embeddings
Zongjie Li, Pingchuan Ma, Huaijin Wang, Shuai Wang, Qiyi Tang, Sen Nie, and Shi Wu · 2022
Closest in time.
An empirical evaluation of github copilot’s code suggestions
Nhan Nguyen and Sarah Nadi · 2022
Closest in time.
A conversational paradigm for program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2022
Closest in time.
Mdpfuzzer: Finding crash-triggering state sequences in models solving the markov decision process
Qi Pang, Yuanyuan Yuan, and Shuai Wang · 2022
Closest in time.
Pop quiz! can a large language model help with reverse engineering?
Hammond Pearce, Benjamin Tan, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri, and Brendan Dolan-Gavitt · 2022
Closest in time.
Pop quiz! can a large language model help with reverse engineering?
Hammond Pearce, Benjamin Tan, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri, and Brendan Dolan-Gavitt · 2022
Closest in time.
Automatic generation of programming exercises and code explanations with large language models
Sami Sarsa, Paul Denny, Arto Hellas, and Juho Leinonen · 2022
Closest in time.
Improving machine translation systems via isotopic replacement
Zeyu Sun, Jie M Zhang, Yingfei Xiong, Mark Harman, Mike Papadakis, and Lu Zhang · 2022
Closest in time.
sem2vec: Semantics-aware assembly tracelet embedding
Huaijin Wang, Pingchuan Ma, Shuai Wang, Qiyi Tang, Sen Nie, and Shi Wu · 2022
Closest in time.
Enhancing dnn-based binary code function search with low-cost equivalence checking
Huaijin Wang, Pingchuan Ma, Yuanyuan Yuan, Zhibo Liu, Shuai Wang, Qiyi Tang, Sen Nie, and Shi Wu · 2022
Closest in time.
Rare tokens degenerate all tokens: Improving neural text generation via adaptive gradient gating for rare token embeddings
Sangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee, Woo-Jong Ryu, and Sungroh Yoon · 2022
Closest in time.
Unveiling hidden dnn defects with decision-based metamorphic testing
Yuanyuan Yuan, Qi Pang, and Shuai Wang · 2022
Closest in time.
Diet code is healthy: Simplifying programs for pre-trained models of code
Zhaowei Zhang, Hongyu Zhang, Beijun Shen, and Xiaodong Gu · 2022
Closest in time.