Fetching the paper…
Reading the bibliography…
Developing modern machine learning (ML) applications is data-centric, of which one fundamental challenge is to understand the influence of data quality to ML training -- "Which training examples are 'guilty' in making the trained ML model predictions inaccurate or unfair?" Modeling data influence for ML training has attracted intensive interest over the last decade, and one popular framework is to compute the Shapley value of each training example with respect to utilities such as validation accuracy and fairness of the trained ML model.
Representation of switching circuits by binary-decision programs
C. Y. Lee. 1959 · 1959
Earlier work this paper cites.
The complexity of computing the permanent
L G Valiant. 1979 · 1979
Earlier work this paper cites.
Graph-based algorithms for boolean function manipulation
Randal E Bryant. 1986 · 1986
Earlier work this paper cites.
A Probabilistic Analysis of the Rocchio Algorithm with TFIDF for Text Categorization
Thorsten Joachims. 1996 · 1996
Earlier work this paper cites.
Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid.. In Kdd , Vol. 96. 202–207
Ron Kohavi et al · 1996
Earlier work this paper cites.
Formal verification using edge-valued binary decision diagrams
Yung-Te Lai, Massoud Pedram, and Sarma B. K. Vrudhula. 1996 · 1996
Earlier work this paper cites.
Algebric decision diagrams and their applications
R Iris Bahar, Erica A Frohm, Charles M Gaona, Gary D Hachtel, Enrico Macii, Abelardo Pardo, and Fabio Somenzi. 1997 · 1997
Earlier work this paper cites.
A survey on knowledge compilation
Marco Cadoli and Francesco M Donini. 1997 · 1997
Earlier work this paper cites.
Generalized hockey stick identities and n-dimensional blockwalking
Peter Ross. 1997 · 1997
Earlier work this paper cites.
Why and where: A characterization of data provenance. In International conference on database theory . Springer, 316–330
Peter Buneman, Sanjeev Khanna, and Tan Wang-Chiew. 2001 · 2001
Earlier work this paper cites.
Affine algebraic decision diagrams (AADDs) and their application to structured probabilistic inference. In IJCAI , Vol. 2005. 1384–1390
Scott Sanner and David McAllester. 2005 · 2005
Earlier work this paper cites.
Provenance semirings. In Proceedings of the twenty-sixth ACM SIGMOD-SIGACT-SIGART symposium on Principles of database systems . 31–40
Todd J Green, Grigoris Karvounarakis, and Val Tannen. 2007 · 2007
Earlier work this paper cites.
Computational complexity: a modern approach
Sanjeev Arora and Boaz Barak. 2009 · 2009
Earlier work this paper cites.
Provenance in databases: Why, how, and where
James Cheney, Laura Chiticariu, and Wang-Chiew Tan. 2009 · 2009
Earlier work this paper cites.
Knowledge compilation meets database theory: Compiling queries to decision diagrams. In ACM International Conference Proceeding Series . 162–173
Abhay Jha and Dan Suciu. 2011 · 2011
Earlier work this paper cites.
Sensitivity analysis and explanations for robust query evaluation in probabilistic databases. In Proceedings of the 2011 international conference on Management of data - SIGMOD ’11 (Athens, Greece). ACM Press, New York, New York, USA
Bhargav Kanagal, Jian Li, and Amol Deshpande. 2011 · 2011
Earlier work this paper cites.
Reverse Data Management
Alexandra Meliou, Wolfgang Gatterbauer, and Dan Suciu. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
Tiresias: The Database Oracle for How-to Queries. In Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data (Scottsdale, Arizona, USA) (SIGMOD ’12) . Association for Computing Machinery, New York, NY, USA, 337–348
Alexandra Meliou and Dan Suciu. 2012 · 2012
Earlier work this paper cites.
Making a Science of Model Search: Hyperparameter Optimization in Hundreds of Dimensions for Vision Architectures
James Bergstra, Daniel Yamins, and David Cox. 2013 · 2013
Earlier work this paper cites.
Searching for exotic particles in high-energy physics with deep learning
Pierre Baldi, Peter Sadowski, and Daniel Whiteson. 2014 · 2014
Earlier work this paper cites.
Causality and Explanations in Databases
Alexandra Meliou, Sudeepa Roy, and Dan Suciu. 2014 · 2014
Earlier work this paper cites.
A Formal Approach to Finding Explanations for Database Queries. In Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data (Snowbird, Utah, USA) (SIGMOD ’14) . Association for Computing Machinery, New York, NY, USA, 1579–1590
Sudeepa Roy and Dan Suciu. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks. In European conference on computer vision . Springer, 818–833
Matthew D Zeiler and Rob Fergus. 2014 · 2014
Cited alongside, same era.
Parallel Feature Selection Inspired by Group Testing
Yingbo Zhou, Utkarsh Porwal, Ce Zhang, Hung Q Ngo, Xuanlong Nguyen, Christopher Ré, and Venu Govindaraju. 2014 · 2014
Cited alongside, same era.
Efficient and robust automated machine learning. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 2 (Montreal, Canada) (NIPS’15) . MIT Press, Cambridge, MA, USA, 2755–2763
Matthias Feurer, Aaron Klein, Katharina Eggensperger, Jost Tobias Springenberg, Manuel Blum, and Frank Hutter. 2015 · 2015
Cited alongside, same era.
Explaining Query Answers with Explanation-Ready Databases
Sudeepa Roy, Laurel Orr, and Dan Suciu. 2015 · 2015
Cited alongside, same era.
Data X-Ray: A Diagnostic Tool for Data Errors. In Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data (Melbourne, Victoria, Australia) (SIGMOD ’15) . Association for Computing Machinery, New York, NY, USA, 1231–1245
Fairness and Machine Learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan. 2019 · 2019
Later among the works it cites.
Data shapley: Equitable valuation of data for machine learning. In International Conference on Machine Learning . PMLR, 2242–2251
Amirata Ghorbani and James Zou. 2019 · 2019
Later among the works it cites.
On the accuracy of influence functions for measuring group effects
Pang Wei W Koh, Kai-Siang Ang, Hubert Teo, and Percy S Liang. 2019 · 2019
Later among the works it cites.
Data Science through the looking glass and what we found there
Fotis Psallidas, Yiwen Zhu, Bojan Karlaš, Matteo Interlandi, Avrilia Floratou, Konstantinos Karanasos, Wentao Wu, Ce Zhang, Subru Krishnan, Carlo Curino, et al · 2019
Later among the works it cites.
SysML: The New Frontier of Machine Learning Systems
Alexander Ratner, Dan Alistarh, Gustavo Alonso, David G Andersen, Peter Bailis, Sarah Bird, Nicholas Carlini, Bryan Catanzaro, Eric Chung, Bill Dally, and Others. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaolan Wang, Xin Luna Dong, and Alexandra Meliou. 2015 · 2015
Cited alongside, same era.
Equality of Opportunity in Supervised Learning. In Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29. Curran Associates, Inc
Moritz Hardt, Eric Price, Eric Price, and Nati Srebro. 2016 · 2016
Cited alongside, same era.
Mllib: Machine learning in apache spark
Xiangrui Meng, Joseph Bradley, Burak Yavuz, Evan Sparks, Shivaram Venkataraman, Davies Liu, Jeremy Freeman, DB Tsai, Manish Amde, Sean Owen, et al · 2016
Cited alongside, same era.
"Why should i trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Neural Architecture Search with Reinforcement Learning
Barret Zoph and Quoc V Le. 2016 · 2016
Cited alongside, same era.
Tfx: A tensorflow-based production-scale machine learning platform. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 1387–1395
Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc, et al · 2017
Cited alongside, same era.
Understanding Black-box Predictions via Influence Functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Cited alongside, same era.
Palm: Machine learning explanations for iterative debugging. In Proceedings of the 2Nd workshop on human-in-the-loop data analytics . 1–6
Sanjay Krishnan and Eugene Wu. 2017 · 2017
Cited alongside, same era.
Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML
Ashvin Agrawal, Rony Chatterjee, Carlo Curino, Avrilia Floratou, Neha Gowdal, Matteo Interlandi, Alekh Jindal, Konstantinos Karanasos, Subru Krishnan, Brian Kroth, et al · 2020
Later among the works it cites.
On Second-Order Group Influence Functions for Black-Box Predictions
Samyadeep Basu, Xuchen You, and Soheil Feizi. 2020 · 2020
Later among the works it cites.
A Unified Architecture for Accelerating Distributed { \{ DNN } \} Training in Heterogeneous { \{ GPU/CPU } \} Clusters. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 463–479
Yimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi, Yong Cui, and Chuanxiong Guo. 2020 · 2020
Later among the works it cites.
Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions
Bojan Karlaš, Peng Li, Renzhi Wu, Nezihe Merve Gürel, Xu Chu, Wentao Wu, and Ce Zhang. 2020 · 2020
Later among the works it cites.
CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification Tasks. In 36th IEEE International Conference on Data Engineering (ICDE 2020)(virtual)
Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang. 2021 · 2020
Later among the works it cites.
PyTorch distributed: experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala. 2020 · 2020
Later among the works it cites.
Distributed Learning Systems with First-Order Methods
Ji Liu, Ce Zhang, and Others. 2020 · 2020
Later among the works it cites.
A Tensor Compiler for Unified Machine Learning Prediction Serving. In 14th { \{ USENIX } \} Symposium on Operating Systems Design and Implementation ( { \{ OSDI } \} 20) . 899–917
Supun Nakandala, Karla Saur, Gyeong-In Yu, Konstantinos Karanasos, Carlo Curino, Markus Weimer, and Matteo Interlandi. 2020 · 2020
Later among the works it cites.
Data Pricing–From Economics to Data Science. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 3553–3554
Jian Pei. 2020 · 2020
Later among the works it cites.
Complaint-driven training data debugging for query 2.0. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data . 1317–1334
Weiyuan Wu, Lampros Flokas, Eugene Wu, and Jiannan Wang. 2020 · 2020
Later among the works it cites.
Explaining by removing: A unified framework for model explanation
Ian Covert, Scott Lundberg, and Su-In Lee. 2021 · 2021
Later among the works it cites.
Retiring Adult: New Datasets For Fair Machine Learning
Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt. 2021 · 2021
Later among the works it cites.
Bagua: scaling up distributed learning with system relaxations
Shaoduo Gan, Jiawei Jiang, Binhang Yuan, Ce Zhang, Xiangru Lian, Rui Wang, Jianbin Chang, Chengjun Liu, Hongmei Shi, Shengzhuo Zhang, Xianghong Li, Tengxu Sun, Sen Yang, and Ji Liu. 2021 · 2021
Later among the works it cites.
Mlinspect: A data distribution debugger for machine learning pipelines. In Proceedings of the 2021 International Conference on Management of Data . 2736–2739
Stefan Grafberger, Shubha Guha, Julia Stoyanovich, and Sebastian Schelter. 2021 · 2021
Later among the works it cites.
Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?
Ruoxi Jia, Xuehui Sun, Jiacen Xu, Ce Zhang, Bo Li, and Dawn Song. 2021 · 2021
Later among the works it cites.
ActiveClean: Interactive data cleaning for statistical modeling
Sanjay Krishnan, Jiannan Wang, Eugene Wu, Michael J Franklin, and Ken Goldberg. [n.d.] · 2021
Later among the works it cites.
Data distribution debugging in machine learning pipelines
Stefan Grafberger, Paul Groth, Julia Stoyanovich, and Sebastian Schelter. 2022 · 2022
Closest in time.
Screening Native ML Pipelines with “ArgusEyes”
Sebastian Schelter, Stefan Grafberger, Shubha Guha, Olivier Sprangers, Bojan Karlaš, and Ce Zhang. 2022 · 2022
Closest in time.