Fetching the paper…
Reading the bibliography…
High-quality data is critical to train performant Machine Learning (ML) models, highlighting the importance of Data Quality Management (DQM).
Efficient task-specific data valuation for nearest neighbor algorithms
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nezihe Merve Gurel, Bo Li, Ce Zhang, Costas J Spanos, and Dawn Song · 1908
Earlier work this paper cites.
Ruoxi Jia, Fan Wu, Xuehui Sun, Jiacen Xu, David Dao, Bhavya Kailkhura, Ce Zhang, Bo Li, and Dawn Song · 1911
Earlier work this paper cites.
Accelerated greedy algorithms for maximizing submodular set functions
Michel Minoux · 1978
Earlier work this paper cites.
The travelling salesman problem: new solvable cases and linkages with the development of approximation algorithms
Fred Glover and Abraham P Punnen · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The enron corpus: A new dataset for email classification research
Bryan Klimt and Yiming Yang · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Wrangler: Interactive visual specification of data transformation scripts
Sean Kandel, Andreas Paepcke, Joseph Hellerstein, and Jeffrey Heer · 2011
Earlier work this paper cites.
Learning submodular functions
Maria-Florina Balcan and Nicholas JA Harvey · 2011
Earlier work this paper cites.
Scaling up biologically-inspired computer vision: A case study in unconstrained face recognition on facebook
Nicolas Pinto, Zak Stone, Todd Zickler, and David Cox · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
Learning fair representations
Rich Zemel, Yu Wu, Kevin Swersky, Toni Pitassi, and Cynthia Dwork · 2013
Earlier work this paper cites.
Classifying spam emails using text and readability features
Rushdi Shams and Robert E Mercer · 2013
Earlier work this paper cites.
Katara: A data cleaning system powered by knowledge bases and crowdsourcing
Xu Chu, John Morcos, Ihab F Ilyas, Mourad Ouzzani, Paolo Papotti, Nan Tang, and Yin Ye · 2015
Earlier work this paper cites.
Addressing the computational issues of the Shapley value with applications in the smart grid
Sasan Maleki · 2015
Earlier work this paper cites.
Lazier than lazy greedy
Baharan Mirzasoleiman, Ashwinkumar Badanidiyuru, Amin Karbasi, Jan Vondrák, and Andreas Krause · 2015
Cited alongside, same era.
Submodularity in data subset selection and active learning
Kai Wei, Rishabh Iyer, and Jeff Bilmes · 2015
Cited alongside, same era.
Activeclean: Interactive data cleaning for statistical modeling
Sanjay Krishnan, Jiannan Wang, Eugene Wu, Michael J Franklin, and Ken Goldberg · 2016
Cited alongside, same era.
Maximization of approximately submodular functions
Thibaut Horel and Yaron Singer · 2016
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Boostclean: Automated error detection and repair for machine learning
Sanjay Krishnan, Michael J Franklin, Ken Goldberg, and Eugene Wu · 2017
Unit testing data with deequ
Sebastian Schelter, Felix Biessmann, Dustin Lange, Tammo Rukat, Phillipp Schmidt, Stephan Seufert, Pierre Brunelle, and Andrey Taptunov · 2019
Later among the works it cites.
Software engineering for machine learning: A case study
Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann · 2019
Later among the works it cites.
Alphaclean: Automatic generation of data cleaning pipelines
Sanjay Krishnan and Eugene Wu · 2019
Later among the works it cites.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao · 2019
Later among the works it cites.
Deepinspect: A black-box trojan detection and mitigation framework for deep neural networks
Huili Chen, Cheng Fu, Jishen Zhao, and Farinaz Koushanfar · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Ruslan Salakhutdinov, and Alexander Smola · 2017
Cited alongside, same era.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2017
Cited alongside, same era.
Trojaning attack on neural networks
Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang · 2017
Cited alongside, same era.
Learning adversarially fair and transferable representations
David Madras, Elliot Creager, Toniann Pitassi, and Richard Zemel · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Sik Kim, Ian EH Yen, and Pradeep Ravikumar · 2018
Cited alongside, same era.
Min Du, Ruoxi Jia, and Dawn Song · 2019
Later among the works it cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Later among the works it cites.
Bojan Karlaš, Peng Li, Renzhi Wu, Nezihe Merve Gürel, Xu Chu, Wentao Wu, and Ce Zhang · 2020
Later among the works it cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan · 2020
Later among the works it cites.
If you like shapley then you’ll love the core, 2020
Tom Yan and Ariel D Procaccia · 2020
Later among the works it cites.
A principled approach to data valuation for federated learning
Tianhao Wang, Johannes Rausch, Ce Zhang, Ruoxi Jia, and Dawn Song · 2020
Later among the works it cites.
On Additive Approximate Submodularity
Flavio Chierichetti, Anirban Dasgupta, and Ravi Kumar · 2020
Later among the works it cites.
Covid-ct-dataset: a ct scan dataset about covid-19
Jinyu Zhao, Yichen Zhang, Xuehai He, and Pengtao Xie · 2020
Later among the works it cites.
Replication-robust payoff-allocation for machine learning data markets
Dongge Han, Michael Wooldridge, Alex Rogers, Shruti Tople, Olga Ohrimenko, and Sebastian Tschiatschek · 2020
Later among the works it cites.
From cleaning before ml to cleaning for ml
Felix Neutatz, Binger Chen, Ziawasch Abedjan, and Eugene Wu · 2021
Closest in time.
Rethinking the backdoor attacks’ triggers: A frequency perspective
Yi Zeng, Won Park, Z Morley Mao, and Ruoxi Jia · 2021
Closest in time.