Fetching the paper…
Reading the bibliography…
This paper provides a comprehensive survey of Machine Learning Testing (ML testing) research.
Sex bias in graduate admissions: Data from Berkeley
Peter J Bickel, Eugene A Hammel, and J William O’Connell · 1975
Earlier work this paper cites.
Symbolic execution and program testing
James C King · 1976
Earlier work this paper cites.
Pseudo-oracles for non-testable programs
Martin D. Davis and Elaine J. Weyuker · 1981
Earlier work this paper cites.
Optimization by simulated annealing
Scott Kirkpatrick, C Daniel Gelatt, and Mario P Vecchi · 1983
Earlier work this paper cites.
Effects of sample size in classifier design
Keinosuke Fukunaga and Raymond R. Hayes · 1989
Earlier work this paper cites.
UCI chess (king-rook vs. king-pawn) data set
Alen Shapiro · 1989
Earlier work this paper cites.
IEEE Std 610.12-1990
Ieee standard glossary of software engineering terminology · 1990
Earlier work this paper cites.
A survey of decision tree classifier methodology
S Rasoul Safavian and David Landgrebe · 1991
Earlier work this paper cites.
The state-based testing of object-oriented programs
Christopher D Turner and David J Robson · 1993
Earlier work this paper cites.
Measuring the vc-dimension of a learning machine
Vladimir Vapnik, Esther Levin, and Yann Le Cun · 1994
Earlier work this paper cites.
An introduction to the bootstrap
Bradley Efron and Robert J Tibshirani · 1994
Earlier work this paper cites.
Software testability: The new verification
Jeffrey M. Voas and Keith W. Miller · 1995
Earlier work this paper cites.
The methodology of n-version programming
Algirdas Avizienis · 1995
Earlier work this paper cites.
A study of cross-validation and bootstrap for accuracy estimation and model selection
Ron Kohavi et al · 1995
Earlier work this paper cites.
Applied linear statistical models
John Neter, Michael H Kutner, Christopher J Nachtsheim, and William Wasserman · 1996
Earlier work this paper cites.
A comparison of event models for naive bayes text classification
Andrew McCallum, Kamal Nigam, et al · 1998
Earlier work this paper cites.
Metamorphic testing: a new approach for generating next test cases
Tsong Y Chen, Shing C Cheung, and Shiu Ming Yiu · 1998
Earlier work this paper cites.
Differential testing for software
William M McKeeman · 1998
Earlier work this paper cites.
Classifier design for computer-aided diagnosis: Effects of finite sample size on the mean performance of classical and neural network classifiers
Heang-Ping Chan, Berkman Sahiner, Robert F Wagner, and Nicholas Petrick · 1999
Earlier work this paper cites.
An empirical evaluation of the mc/dc coverage criterion on the hete-2 satellite software
Arnaud Dupuy and Nancy Leveson · 2000
Earlier work this paper cites.
Feature selection and classifier performance in computer-aided diagnosis: The effect of finite sample size
Berkman Sahiner, Heang-Ping Chan, Nicholas Petrick, Robert F Wagner, and Lubomir Hadjiiski · 2000
Earlier work this paper cites.
Data cleaning: Problems and current approaches
Erhard Rahm and Hong Hai Do · 2000
Earlier work this paper cites.
Search-based software engineering
Mark Harman and Bryan F Jones · 2001
Earlier work this paper cites.
Gui testing: Pitfalls and process
Atif M Memon · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Kernel methods for pattern analysis
John Shawe-Taylor, Nello Cristianini, et al · 2004
Earlier work this paper cites.
Feedback loop: The civil rights act of 1964 and its progeny
Drew S Days III · 2004
Earlier work this paper cites.
Search-based software test data generation: a survey
Phil McMinn · 2004
Earlier work this paper cites.
The problem of overfitting
Douglas M Hawkins · 2004
Earlier work this paper cites.
Evaluation of different fitness functions for the evolutionary testing of an autonomous parking system
Joachim Wegener and Oliver Bühler · 2004
Earlier work this paper cites.
Cute: a concolic unit testing engine for c
Koushik Sen, Darko Marinov, and Gul Agha · 2005
Earlier work this paper cites.
Software and hardware testing using combinatorial covering suites
Alan Hartman · 2005
Earlier work this paper cites.
Mujava: an automated class mutation system
Yu-Seung Ma, Jeff Offutt, and Yong Rae Kwon · 2005
Earlier work this paper cites.
Why question machine learning evaluation methods
Nathalie Japkowicz · 2006
Earlier work this paper cites.
An approach to software testing of machine learning applications
Christian Murphy, Gail E Kaiser, and Marta Arias · 2007
Earlier work this paper cites.
The rademacher complexity of co-regularized kernel classes
David S Rosenberg and Peter L Bartlett · 2007
Earlier work this paper cites.
A multi-objective approach to search-based test data generation
Kiran Lakhotia, Mark Harman, and Phil McMinn · 2007
Earlier work this paper cites.
Parameterizing Random Test Data According to Equivalence Classes
Christian Murphy, Gail Kaiser, and Marta Arias · 2007
Earlier work this paper cites.
Concentration inequalities and model selection
Pascal Massart · 2007
Earlier work this paper cites.
Report to the congress on credit scoring and its effects on the availability and affordability of credit
US Federal Reserve · 2007
Earlier work this paper cites.
Constructing subtle faults using higher order mutation testing (best paper award winner)
Yue Jia and Mark Harman · 2008
Earlier work this paper cites.
Properties of machine learning applications for use in metamorphic testing
Christian Murphy, Gail E. Kaiser, Lifeng Hu, and Leon Wu · 2008
Earlier work this paper cites.
Fairness analysis in requirements assignments
Anthony Finkelstein, Mark Harman, Afshin Mansouri, Jian Ren, and Yuanyuan Zhang · 2008
Earlier work this paper cites.
Isolation forest
Fei Tony Liu, Kai Ming Ting, and Zhi-Hua Zhou · 2008
Earlier work this paper cites.
A systematic review of search-based testing for non-functional system properties
Wasif Afzal, Richard Torkar, and Robert Feldt · 2009
Earlier work this paper cites.
Using JML Runtime Assertion Checking to Automate Metamorphic Testing in Applications Without Test Oracles
Christian Murphy, Kuang Shen, and Gail Kaiser · 2009
Earlier work this paper cites.
Application of metamorphic testing to supervised classifiers
Xiaoyuan Xie, Joshua Ho, Christian Murphy, Gail Kaiser, Baowen Xu, and Tsong Yueh Chen · 2009
Earlier work this paper cites.
The weka data mining software: an update
Mark Hall, Eibe Frank, Geoffrey Holmes, Bernhard Pfahringer, Peter Reutemann, and Ian H Witten · 2009
Earlier work this paper cites.
Automatic System Testing of Programs Without Test Oracles
Christian Murphy, Kuang Shen, and Gail Kaiser · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
A search based approach to fairness analysis in requirements assignments to aid negotiation, mediation and decision making
Anthony Finkelstein, Mark Harman, Afshin Mansouri, Jian Ren, and Yuanyuan Zhang · 2009
Earlier work this paper cites.
A theoretical and empirical study of search-based testing: Local, global, and hybrid search
Mark Harman and Phil McMinn · 2010
Earlier work this paper cites.
IEEE Std 1044-2009 (Revision of IEEE Std 1044-1993)
Ieee standard classification for software anomalies · 2010
Earlier work this paper cites.
Test generation via dynamic symbolic execution for mutation testing
Lingming Zhang, Tao Xie, Lu Zhang, Nikolai Tillmann, Jonathan De Halleux, and Hong Mei · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Unit-Testing Statistical Software
Maciej Pacula · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Differential privacy
Cynthia Dwork · 2011
Earlier work this paper cites.
An analysis and survey of the development of mutation testing
Yue Jia and Mark Harman · 2011
Earlier work this paper cites.
Testing and validating machine learning classifiers by metamorphic testing
Xiaoyuan Xie, Joshua W. K. Ho, Christian Murphy, Gail Kaiser, Baowen Xu, and Tsong Yueh Chen · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Squeeziness: An information theoretic measure for avoiding fault masking
David Clark and Robert M. Hierons · 2012
Earlier work this paper cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Earlier work this paper cites.
Non-functional requirements in software engineering
Lawrence Chung, Brian A Nixon, Eric Yu, and John Mylopoulos · 2012
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel · 2012
Earlier work this paper cites.
An empirical study of bugs in machine learning systems
Ferdian Thung, Shaowei Wang, David Lo, and Lingxiao Jiang · 2012
Earlier work this paper cites.
Classifier variability: accounting for training and testing
Weijie Chen, Brandon D Gallas, and Waleed A Yousef · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
A systematic review of software robustness
Ali Shahrokni and Robert Feldt · 2013
Earlier work this paper cites.
State of the art: Dynamic symbolic execution for automated test generation
Ting Chen, Xiao-song Zhang, Shi-ze Guo, Hong-yuan Li, and Yue Wu · 2013
Earlier work this paper cites.
On the assessment of the added value of new predictive biomarkers
Weijie Chen, Frank W Samuelson, Brandon D Gallas, Le Kang, Berkman Sahiner, and Nicholas Petrick · 2013
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
A Geiger, P Lenz, C Stiller, and R Urtasun · 2013
Earlier work this paper cites.
http://contagiodump.blogspot.com/2013/03/16800-clean-and-11960-malicious-files.html
contagio malware dump · 2013
Earlier work this paper cites.
An analysis of the relationship between conditional entropy and failed error propagation in software testing
Kelly Androutsopoulos, David Clark, Haitao Dan, Mark Harman, and Robert Hierons · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Earlier work this paper cites.
Thoughtful machine learning: A test-driven approach
Matthew Kirk · 2014
Earlier work this paper cites.
Search-based inference of polynomial metamorphic relations
Jie Zhang, Junjie Chen, Dan Hao, Yingfei Xiong, Bing Xie, Lu Zhang, and Hong Mei · 2014
Earlier work this paper cites.
Guidelines for snowballing in systematic literature studies and a replication in software engineering
Claes Wohlin · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Compiler validation via equivalence modulo inputs
Vu Le, Mehrdad Afshari, and Zhendong Su · 2014
Earlier work this paper cites.
Unit tests for stochastic optimization
Tom Schaul, Ioannis Antonoglou, and David Silver · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
A data-driven approach to predict the success of bank telemarketing
Sérgio Moro, Paulo Cortez, and Paulo Rita · 2014
Earlier work this paper cites.
Drebin: Effective and explainable detection of android malware in your pocket
Daniel Arp, Michael Spreitzenbarth, Malte Hubner, Hugo Gascon, Konrad Rieck, and CERT Siemens · 2014
Earlier work this paper cites.
Deepdriving: Learning affordance for direct perception in autonomous driving
Chenyi Chen, Ari Seff, Alain Kornhauser, and Jianxiong Xiao · 2015
Earlier work this paper cites.
The oracle problem in software testing: A survey
Earl T Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo · 2015
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Earlier work this paper cites.
Machine learning: Trends, perspectives, and prospects
Michael I Jordan and Tom M Mitchell · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Lifelong machine learning test
Lianghao Li and Qiang Yang · 2015
Cited alongside, same era.
Fairness constraints: Mechanisms for fair classification
Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez Rodriguez, and Krishna P Gummadi · 2015
Cited alongside, same era.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei · 2015
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning · 2015
Cited alongside, same era.
Calibrating probability with undersampling for unbalanced classification
Andrea Dal Pozzolo, Olivier Caelen, Reid A Johnson, and Gianluca Bontempi · 2015
Cited alongside, same era.
Simulation-based adversarial test generation for autonomous vehicles with machine learning components
Cumhur Erkan Tuncali, Georgios Fainekos, Hisahiro Ito, and James Kapinski · 2018
Later among the works it cites.
Symbolic execution for deep neural networks
Divya Gopinath, Kaiyuan Wang, Mengshi Zhang, Corina S Pasareanu, and Sarfraz Khurshid · 2018
Later among the works it cites.
Automated test generation to detect individual discrimination in ai models
Aniya Agarwal, Pranay Lohia, Seema Nagar, Kuntal Dey, and Diptikalyan Saha · 2018
Later among the works it cites.
Concolic testing for deep neural networks
Youcheng Sun, Min Wu, Wenjie Ruan, Xiaowei Huang, Marta Kwiatkowska, and Daniel Kroening · 2018
Later among the works it cites.
Identifying implementation bugs in machine learning based image classifiers using metamorphic testing
Anurag Dwarakanath, Manish Ahuja, Samarth Sikand, Raghotham M. Rao, R. P. Jagadeesh Chandra Bose, Neville Dubash, and Sanjay Podder · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Introduction to software testing
Paul Ammann and Jeff Offutt · 2016
Cited alongside, same era.
Integrating symbolic and statistical methods for testing intelligent systems: Applications to machine learning and computer vision
A. Ramanathan, L. L. Pullum, F. Hussain, D. Chakrabarty, and S. K. Jha · 2016
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Cited alongside, same era.
Data poisoning attacks against autoregressive models
Scott Alfeld, Xiaojin Zhu, and Paul Barford · 2016
Cited alongside, same era.
The mythos of model interpretability
Zachary C Lipton · 2016
Cited alongside, same era.
European union regulations on algorithmic decision-making and a" right to explanation"
Bryce Goodman and Seth Flaxman · 2016
Cited alongside, same era.
A framework for ensuring the quality of a big data service
Junhua Ding, Dongmei Zhang, and Xin-Hua Hu · 2016
Cited alongside, same era.
Later among the works it cites.
Failing to learn: Autonomously identifying perception failures for self-driving cars
Manikandasriram Srinivasan Ramanagopal, Cyrus Anderson, Ram Vasudevan, and Matthew Johnson-Roberson · 2018
Later among the works it cites.
Mettle: A metamorphic testing approach to validating unsupervised machine learning methods, 2018
Xiaoyuan Xie, Zhiyi Zhang, Tsong Yueh Chen, Yang Liu, Pak-Lok Poon, and Baowen Xu · 2018
Later among the works it cites.
Dataset diversity for metamorphic testing of machine learning software
Shin Nakajima · 2018
Later among the works it cites.
Guiding deep learning system testing using surprise adequacy
Jinhan Kim, Robert Feldt, and Shin Yoo · 2018
Later among the works it cites.
Metamorphic Testing for Machine Translations: MT4MT
L. Sun and Z. Q. Zhou · 2018
Later among the works it cites.
A monte carlo method for metamorphic testing of machine translation services
Daniel Pesu, Zhi Quan Zhou, Jingfeng Zhen, and Dave Towey · 2018
Later among the works it cites.
Multiple-implementation testing of supervised learning software
Siwakorn Srisakaokul, Zhengkai Wu, Angello Astorga, Oreoluwa Alebiosu, and Tao Xie · 2018
Later among the works it cites.
Syneva: Evaluating ml programs by mirror program synthesis
Yi Qin, Huiyan Wang, Chang Xu, Xiaoxing Ma, and Jian Lu · 2018
Later among the works it cites.
Model assertions for debugging machine learning
Daniel Kang, Deepti Raghavan, Peter Bailis, and Matei Zaharia · 2018
Later among the works it cites.
Youcheng Sun, Xiaowei Huang, and Daniel Kroening · 2018
Later among the works it cites.
Combinatorial testing for deep learning systems
Lei Ma, Fuyuan Zhang, Minhui Xue, Bo Li, Yang Liu, Jianjun Zhao, and Yadong Wang · 2018
Later among the works it cites.
DeepMutation: Mutation Testing of Deep Learning Systems, 2018
Lei Ma, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Felix Juefei-Xu, Chao Xie, Li Li, Yang Liu, Jianjun Zhao, and Yadong Wang · 2018
Later among the works it cites.
MuNN: Mutation Analysis of Neural Networks
Weijun Shen, Jun Wan, and Zhenyu Chen · 2018
Later among the works it cites.
An empirical study on tensorflow program bugs
Yuhao Zhang, Yifan Chen, Shing-Chi Cheung, Yingfei Xiong, and Lu Zhang · 2018
Later among the works it cites.
Hands off the wheel in autonomous vehicles?: A systems perspective on over a million miles of field data
S. S. Banerjee, S. Jha, J. Cyriac, Z. T. Kalbarczyk, and R. K. Iyer · 2018
Later among the works it cites.
MODE: Automated Neural Network Model Debugging via State Differential Analysis and Input Selection
Shiqing Ma, Yingqi Liu, Wen-Chuan Lee, Xiangyu Zhang, and Ananth Grama · 2018
Later among the works it cites.
Mistique: A system to store and query model intermediates for model diagnosis
Manasi Vartak, Joana M F da Trindade, Samuel Madden, and Matei Zaharia · 2018
Later among the works it cites.
Telemade: A testing framework for learning-based malware detection systems
Wei Yang and Tao Xie · 2018
Later among the works it cites.
A test architecture for machine learning product
Yasuharu Nishi, Satoshi Masuda, Hideto Ogawa, and Keiji Uetsuki · 2018
Later among the works it cites.
Automating large-scale data quality verification
Sebastian Schelter, Dustin Lange, Philipp Schmidt, Meltem Celikel, Felix Biessmann, and Andreas Grafberger · 2018
Later among the works it cites.
Test data reuse for evaluation of adaptive machine learning algorithms: over-fitting to a fixed’test’dataset and a potential solution
Alexej Gossmann, Aria Pezeshk, and Berkman Sahiner · 2018
Later among the works it cites.
Global robustness evaluation of deep neural networks with provable guarantees for l0 norm
Wenjie Ruan, Min Wu, Youcheng Sun, Xiaowei Huang, Daniel Kroening, and Marta Kwiatkowska · 2018
Later among the works it cites.
Deepsafe: A data-driven approach for assessing robustness of neural networks
Divya Gopinath, Guy Katz, Corina S Păsăreanu, and Clark Barrett · 2018
Later among the works it cites.
Technical report on the cleverhans v2.1.0 adversarial examples library
Nicolas Papernot, Fartash Faghri, Nicholas Carlini, Ian Goodfellow, Reuben Feinman, Alexey Kurakin, Cihang Xie, Yash Sharma, Tom Brown, Aurko Roy, Alexander Matyasko, Vahid Behzadan, Karen Hambardzumyan, Zhishuai Zhang, Yi-Lin Juang, Zhi Li, Ryan Sheatsley, Abhibhav Garg, Jonathan Uesato, Willi Gierke, Yinpeng Dong, David Berthelot, Paul Hendricks, Jonas Rauber, and Rujun Long · 2018
Later among the works it cites.
AVFI: Fault injection for autonomous vehicles
Saurabh Jha, Subho S Banerjee, James Cyriac, Zbigniew T Kalbarczyk, and Ravishankar K Iyer · 2018
Later among the works it cites.
Kayotee: A fault injection-based system to assess the safety and reliability of autonomous vehicles to faults and errors
Saurabh Jha, Timothy Tsai, Siva Hari, Michael Sullivan, Zbigniew Kalbarczyk, Stephen W Keckler, and Ravishankar K Iyer · 2018
Later among the works it cites.
Fairness definitions explained
Sahil Verma and Julia Rubin · 2018
Later among the works it cites.
Nripsuta Saxena, Karen Huang, Evan DeFilippis, Goran Radanovic, David Parkes, and Yang Liu · 2018
Later among the works it cites.
A reductions approach to fair classification
Alekh Agarwal, Alina Beygelzimer, Miroslav Dudík, John Langford, and Hanna Wallach · 2018
Later among the works it cites.
Themis: Automatically testing software for discrimination
Rico Angell, Brittany Johnson, Yuriy Brun, and Alexandra Meliou · 2018
Later among the works it cites.
Causal testing: Finding defects’ root causes
Brittany Johnson, Yuriy Brun, and Alexandra Meliou · 2018
Later among the works it cites.
Metamorphic relations for enhancing system understanding and use
Zhi Quan Zhou, Liqun Sun, Tsong Yueh Chen, and Dave Towey · 2018
Later among the works it cites.
Calibration of medical diagnostic classifier scores to the probability of disease
Weijie Chen, Berkman Sahiner, Frank Samuelson, Aria Pezeshk, and Nicholas Petrick · 2018
Later among the works it cites.
Detecting violations of differential privacy
Zeyu Ding, Yuxin Wang, Guanhong Wang, Danfeng Zhang, and Daniel Kifer · 2018
Later among the works it cites.
Dp-finder: Finding differential privacy violations by sampling and optimization
Benjamin Bichsel, Timon Gehr, Dana Drachsler-Cohen, Petar Tsankov, and Martin Vechev · 2018
Later among the works it cites.
Detecting adversarial samples for deep neural networks through mutation testing
Jingyi Wang, Jun Sun, Peixin Zhang, and Xinyu Wang · 2018
Later among the works it cites.
An orchestrated empirical study on deep learning frameworks and platforms
An Orchestrated Empirical Study on Deep Learning Frameworks and Lei Ma Qiang Hu Ruitao Feng Li Li Yang Liu Jianjun Zhao Xiaohong Li Platforms Qianyu Guo, Xiaofei Xie · 2018
Later among the works it cites.
Adaptation of general concepts of software testing to neural networks
Yu. L. Karpov, L. E. Karpov, and Yu. G. Smetanin · 2018
Later among the works it cites.
Neural-machine-translation-based commit message generation: How far are we?
Zhongxin Liu, Xin Xia, Ahmed E. Hassan, David Lo, Zhenchang Xing, and Xinyu Wang · 2018
Later among the works it cites.
Ariadne: Analysis for machine learning programs
Julian Dolby, Avraham Shinnar, Allison Allain, and Jenna Reinen · 2018
Later among the works it cites.
Security risks in deep learning implementations
Qixue Xiao, Kang Li, Deyue Zhang, and Weilin Xu · 2018
Later among the works it cites.
Manifesting bugs in machine learning code: An explorative study with mutation testing
Dawei Cheng, Chun Cao, Chang Xu, and Xiaoxing Ma · 2018
Later among the works it cites.
Testing vision-based control systems using learnable evolutionary algorithms
Raja Ben Abdessalem, Shiva Nejati, Lionel C Briand, and Thomas Stifter · 2018
Later among the works it cites.
Testing autonomous cars for feature interaction failures using many-objective search
Raja Ben Abdessalem, Annibale Panichella, Shiva Nejati, Lionel C Briand, and Thomas Stifter · 2018
Later among the works it cites.
Testing untestable neural machine translation: An industrial case
Wujie Zheng, Wenyu Wang, Dian Liu, Changrong Zhang, Qinsong Zeng, Yuetang Deng, Wei Yang, Pinjia He, and Tao Xie · 2018
Later among the works it cites.
Fruit recognition from images using deep learning
Horea Mure · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Later among the works it cites.
From start-ups to scale-ups: opportunities and open problems for static and dynamic program analysis
Mark Harman and Peter O’Hearn · 2018
Later among the works it cites.
Quality assurance of machine learning software
Shin NAKAJIMA · 2018
Later among the works it cites.
Software engineering for machine learning: A case study
Saleema Amershi, Andrew Begel, Christian Bird, Rob DeLine, Harald Gall, Ece Kamar, Nachi Nagappan, Besmira Nushi, and Tom Zimmermann · 2019
Closest in time.
Detecting overfitting via adversarial examples
Roman Werpachowski, András György, and Csaba Szepesvári · 2019
Closest in time.
Data Validation for Machine Learning
Eric Breck, Neoklis Polyzotis, Sudip Roy, Steven Whang, and Martin Zinkevich · 2019
Closest in time.
Cradle: Cross-backend validation to detect and localize bugs in deep learning libraries
Weizhen Qi Lin Tan Viet Hung Pham, Thibaud Lutellier · 2019
Closest in time.
Perturbed Model Validation: A New Framework to Validate Model Relevance
Jie Zhang, Earl T Barr, Benjamin Guedj, Mark Harman, and John Shawe-Taylor · 2019
Closest in time.
Deepbase: Deep inspection of neural networks
Thibault Sellam, Kevin Lin, Ian Yiran Huang, Michelle Yang, Carl Vondrick, and Eugene Wu · 2019
Closest in time.
Interpretable Machine Learning
Christoph Molnar · 2019
Closest in time.
Fairness-aware programming
Aws Albarghouthi and Samuel Vinitsky · 2019
Closest in time.
Testing neural program analyzers, 2019
Md Rafiqul Islam Rabin, Ke Wang, and Mohammad Amin Alipour · 2019
Closest in time.
code2vec: Learning distributed representations of code
Uri Alon, Meital Zilberstein, Omer Levy, and Eran Yahav · 2019
Closest in time.
Rigorous agent evaluation: An adversarial approach to uncover catastrophic failures
Jonathan Uesato, Ananya Kumar, Csaba Szepesvari, Tom Erez, Avraham Ruderman, Keith Anderson, Nicolas Heess, Pushmeet Kohli, et al · 2019
Closest in time.
Metamorphic testing of driverless cars
Zhi Quan Zhou and Liqun Sun · 2019
Closest in time.
Ml-based fault injection for autonomous vehicles
Saurabh Jha, Subho S Banerjee, Timothy Tsai, Siva KS Hari, Michael B Sullivan, Zbigniew T Kalbarczyk, Stephen W Keckler, and Ravishankar K Iyer · 2019
Closest in time.
Grammar based directed testing of machine learning systems
Sakshi Udeshi and Sudipta Chattopadhyay · 2019
Closest in time.
Testing machine learning algorithms for balanced data usage
Arnab Sharma and Heike Wehrheim · 2019
Closest in time.
A study of oracle approximations in testing deep learning libraries
Mahdi Nejadgholi and Jinqiu Yang · 2019
Closest in time.
Towards improved testing for deep learning
Jasmine Sekhon and Cody Fleming · 2019
Closest in time.
Deepct: Tomographic combinatorial testing for deep learning systems
Lei Ma, Felix Juefei-Xu, Minhui Xue, Bo Li, Li Li, Yang Liu, and Jianjun Zhao · 2019
Closest in time.
Structural coverage criteria for neural networks could be misleading
Zenan Li, Xiaoxing Ma, Chang Xu, and Chun Cao · 2019
Closest in time.
Input prioritization for testing neural networks
Taejoon Byun, Vaibhav Sharma, Abhishek Vijayakumar, Sanjai Rayadurgam, and Darren Cofer · 2019
Closest in time.
A noise-sensitivity-analysis-based test prioritization technique for deep neural networks
Long Zhang, Xuechao Sun, Yong Li, Zhenyu Zhang, and Yang Feng · 2019
Closest in time.
Boosting Operational DNN Testing Efficiency through Conditioning
Zenan Li, Xiaoxing Ma, Chang Xu, Chun Cao, Jingwei Xu, and Jian Lu · 2019
Closest in time.
Test selection for deep learning systems, 2019
Wei Ma, Mike Papadakis, Anestis Tsakmalis, Maxime Cordy, and Yves Le Traon · 2019
Closest in time.
Storm: Program reduction for testing and debugging probabilistic programming systems
Saikat Dutta, Wenxian Zhang, Zixin Huang, and Sasa Misailovic · 2019
Closest in time.
Preventing undesirable behavior of intelligent machines
Philip S Thomas, Bruno Castro da Silva, Andrew G Barto, Stephen Giguere, Yuriy Brun, and Emma Brunskill · 2019
Closest in time.
Robustness of neural networks: A probabilistic and practical perspective
Ravi Mangal, Aditya Nori, and Alessandro Orso · 2019
Closest in time.
Towards a bayesian approach for assessing faulttolerance of deep neural networks
Subho S Banerjee, James Cyriac, Saurabh Jha, Zbigniew T Kalbarczyk, and Ravishankar K Iyer · 2019
Closest in time.
Towards testing of deep learning systems with training set reduction
Helge Spieker and Arnaud Gotlieb · 2019
Closest in time.
Offline Contextual Bandits with High Probability Fairness Guarantees
Blossom Metevier, Stephen Giguere, Sarah Brockman, Ari Kobren, Yuriy Brun, Emma Brunskill, and Philip Thomas · 2019
Closest in time.
Assessing the Local Interpretability of Machine Learning Models
Sorelle A. Friedler, Chitradeep Dutta Roy, Carlos Scheidegger, and Dylan Slack · 2019
Closest in time.
Adversarial sample detection for deep neural network through model mutation testing
Jingyi Wang, Guoliang Dong, Jun Sun, Xinyu Wang, and Peixin Zhang · 2019
Closest in time.
Alphaclean: automatic generation of data cleaning pipelines
Sanjay Krishnan and Eugene Wu · 2019
Closest in time.
Open questions in testing of learned computer vision functions for automated driving
Matthias Woehrle, Christoph Gladisch, and Christian Heinzemann · 2019
Closest in time.
Testing untestable neural machine translation: An industrial case
Wujie Zheng, Wenyu Wang, Dian Liu, Changrong Zhang, Qinsong Zeng, Yuetang Deng, Wei Yang, Pinjia He, and Tao Xie · 2019
Closest in time.
Detecting failures of neural machine translation in the absence of reference translations
Wenyu Wang, Wujie Zheng, Dian Liu, Changrong Zhang, Qinsong Zeng, Yuetang Deng, Wei Yang, Pinjia He, and Tao Xie · 2019
Closest in time.
Do pseudo test suites lead to inflated correlation in measuring test effectiveness?
Jie Zhang, Lingming Zhang, Dan Hao, Meng Wang, and Lu Zhang · 2019
Closest in time.
Automatic testing and improvement of machine translation
Zeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis, and Lu Zhang · 2020
Closest in time.