Fetching the paper…
Reading the bibliography…
Data-centric AI is at the center of a fundamental shift in software engineering where machine learning becomes the new software, powered by big data and computing infrastructure.
A survey of sampling from contaminated distributions
John W Tukey · 1960
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
Unsupervised word sense disambiguation rivaling supervised methods
David Yarowsky · 1995
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
Avrim Blum and Tom Mitchell · 1998
Earlier work this paper cites.
A survey of outlier detection methodologies
Victoria J. Hodge and Jim Austin · 2004
Earlier work this paper cites.
Democratic co-learning
Yan Zhou and Sally A. Goldman · 2004
Earlier work this paper cites.
Tri-training: Exploiting unlabeled data using three classifiers
Zhi-Hua Zhou and Ming Li · 2005
Earlier work this paper cites.
Semi-Supervised Learning Literature Survey
Xiaojin Zhu · 2005
Earlier work this paper cites.
Casting out demons: Sanitizing training data for anomaly sensors
Gabriela F. Cretu, Angelos Stavrou, Michael E. Locasto, Salvatore J. Stolfo, and Angelos D. Keromytis · 2008
Earlier work this paper cites.
Alpha-investing: a procedure for sequential control of expected false discoveries
Dean P. Foster and Robert A. Stine · 2008
Earlier work this paper cites.
Denoising natural images based on a modified sparse coding algorithm
Li Shang · 2008
Earlier work this paper cites.
Get another label? improving data quality and data mining using multiple, noisy labelers
Victor S. Sheng, Foster J. Provost, and Panagiotis G. Ipeirotis · 2008
Earlier work this paper cites.
Distant supervision for relation extraction without labeled data
Mike Mintz, Steven Bills, Rion Snow, and Daniel Jurafsky · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2010
Earlier work this paper cites.
Data preprocessing techniques for classification without discrimination
Faisal Kamiran and Toon Calders · 2011
Earlier work this paper cites.
mice: Multivariate imputation by chained equations in r
Stef van Buuren and Karin Groothuis-Oudshoorn · 2011
Earlier work this paper cites.
Principles of Data Integration
AnHai Doan, Alon Y. Halevy, and Zachary G. Ives · 2012
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard S. Zemel · 2012
Earlier work this paper cites.
Fairness-aware classifier with prejudice remover regularizer
Toshihiro Kamishima, Shotaro Akaho, Hideki Asoh, and Jun Sakuma · 2012
Earlier work this paper cites.
Active Learning
Burr Settles · 2012
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli · 2013
Earlier work this paper cites.
Hermoupolis: A trajectory generator for simulating generalized mobility patterns
Nikos Pelekis, Christos Ntrigkogias, Panagiotis Tampakis, Stylianos Sideridis, and Yannis Theodoridis · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Certifying and removing disparate impact
Michael Feldman, Sorelle A. Friedler, John Moeller, Carlos Scheidegger, and Suresh Venkatasubramanian · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Training deep neural networks on noisy labels with bootstrapping
Scott E. Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Recommender Systems Handbook
Francesco Ricci, Lior Rokach, and Bracha Shapira, editors · 2015
Earlier work this paper cites.
Hidden technical debt in machine learning systems
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Dennison · 2015
Earlier work this paper cites.
Data wrangling: The challenging yourney from the wild to the lake
Ignacio G. Terrizzano, Peter M. Schwarz, Mary Roth, and John E. Colino · 2015
Earlier work this paper cites.
Self-labeled techniques for semi-supervised learning: taxonomy, software and empirical study
Isaac Triguero, Salvador García, and Francisco Herrera · 2015
Earlier work this paper cites.
SEEDB: Efficient data-driven visualization recommendations to support visual analytics
Manasi Vartak, Sajjadur Rahman, Samuel Madden, Aditya G. Parameswaran, and Neoklis Polyzotis · 2015
Earlier work this paper cites.
Machine bias: There’s software used across the country to predict future criminals. And its biased against blacks., 2016
J. Angwin, J. Larson, S. Mattu, and L. Kirchner · 2016
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
Compas risk scales: Demonstrating accuracy equity and predictive parity
William Dieterich, Christina Mendoza, and Tim Brennan · 2016
Earlier work this paper cites.
Goods: Organizing Google’s datasets
Alon Y. Halevy, Flip Korn, Natalya Fridman Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy, and Steven Euijong Whang · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro · 2016
Earlier work this paper cites.
Activeclean: Interactive data cleaning for statistical modeling
Sanjay Krishnan, Jiannan Wang, Eugene Wu, Michael J. Franklin, and Ken Goldberg · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Distillation as a defense to adversarial perturbations against deep neural networks
Nicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami · 2016
Earlier work this paper cites.
TFX: A tensorflow-based production-scale machine learning platform
Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Xin Zhang, and Martin Zinkevich · 2017
Earlier work this paper cites.
Fairness in criminal justice risk assessments: The state of the art, 2017
Richard Berk, Hoda Heidari, Shahin Jabbari, Michael Kearns, and Aaron Roth · 2017
Earlier work this paper cites.
Query optimization for dynamic imputation
José Cambronero, John K. Feser, Micah J. Smith, and Samuel Madden · 2017
Earlier work this paper cites.
Active bias: Training more accurate neural networks by emphasizing high variance samples
Haw-Shiuan Chang, Erik G. Learned-Miller, and Andrew McCallum · 2017
Earlier work this paper cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Alexandra Chouldechova · 2017
Earlier work this paper cites.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Earlier work this paper cites.
NIPS 2016 tutorial: Generative adversarial networks
Ian J. Goodfellow · 2017
Earlier work this paper cites.
Meet michelangelo: Uber’s machine learning platform, 2017
M. Hermann, J.and Del Baso · 2017
Earlier work this paper cites.
Avoiding discrimination through causal reasoning
Niki Kilbertus, Mateo Rojas-Carulla, Giambattista Parascandolo, Moritz Hardt, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva · 2017
Earlier work this paper cites.
Decoupling ”when to update” from ”how to update”
Eran Malach and Shai Shalev-Shwartz · 2017
Earlier work this paper cites.
Magnet: A two-pronged defense against adversarial examples
Dongyu Meng and Hao Chen · 2017
Earlier work this paper cites.
On detecting adversarial perturbations
Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff · 2017
Earlier work this paper cites.
Making deep neural networks robust to label noise: A loss correction approach
Giorgio Patrini, Alessandro Rozza, Aditya Krishna Menon, Richard Nock, and Lizhen Qu · 2017
Earlier work this paper cites.
On fairness and calibration
Geoff Pleiss, Manish Raghavan, Felix Wu, Jon M. Kleinberg, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Data management challenges in production machine learning
Neoklis Polyzotis, Sudip Roy, Steven Euijong Whang, and Martin Zinkevich · 2017
Earlier work this paper cites.
Snorkel: Rapid training data creation with weak supervision
Alexander Ratner, Stephen H. Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré · 2017
Earlier work this paper cites.
Learning to compose domain-specific transformations for data augmentation
Alexander J. Ratner, Henry R. Ehrenberg, Zeshan Hussain, Jared Dunnmon, and Christopher Ré · 2017
Earlier work this paper cites.
Holoclean: Holistic data repairs with probabilistic inference
Theodoros Rekatsinas, Xu Chu, Ihab F. Ilyas, and Christopher Ré · 2017
Earlier work this paper cites.
Automatically tracking metadata and provenance of machine learning experiments
S. Schelter, Joos-Hendrik Böse, Johannes Kirschnick, T. Klein, and Stephan Seufert · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Earlier work this paper cites.
Fairness beyond disparate treatment & disparate impact: Learning classification without disparate mistreatment
Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez-Rodriguez, and Krishna P. Gummadi · 2017
Cited alongside, same era.
Fairness constraints: Mechanisms for fair classification
Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez-Rodriguez, and Krishna P. Gummadi · 2017
Cited alongside, same era.
Controlling false discoveries during interactive data exploration
Zheguang Zhao, Lorenzo De Stefani, Emanuel Zgraggen, Carsten Binnig, Eli Upfal, and Tim Kraska · 2017
Cited alongside, same era.
A reductions approach to fair classification
Alekh Agarwal, Alina Beygelzimer, Miroslav Dudík, John Langford, and Hanna M. Wallach · 2018
Cited alongside, same era.
Ten years of webtables
Michael J. Cafarella, Alon Y. Halevy, Hongrae Lee, Jayant Madhavan, Cong Yu, Daisy Zhe Wang, and Eugene Wu · 2018
Cited alongside, same era.
Machine learning and big data: What is important?
Michael Stonebraker and El Kindi Rezig · 2019
Later among the works it cites.
Algorithmic fairness: Measures, methods and representations
Suresh Venkatasubramanian · 2019
Later among the works it cites.
Learning with noisy labels for sentence-level sentiment classification
Hao Wang, Bing Liu, Chaozhuo Li, Yan Yang, and Tianrui Li · 2019
Later among the works it cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh, Youngjoon Yoo, and Junsuk Choe · 2019
Later among the works it cites.
Transferable clean-label poisoning attacks on deep neural nets
Chen Zhu, W. Ronny Huang, Hengduo Li, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2019
Later among the works it cites.
https://www.datanami.com/2020/07/06/data-prep-still-dominates-data-scientists-time-survey-finds/
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay · 2018
Cited alongside, same era.
Recurrent neural networks for multivariate time series with missing values
Zhengping Che, Sanjay Purushotham, Kyunghyun Cho, David Sontag, and Yan Liu · 2018
Cited alongside, same era.
Why is my classifier discriminatory?
Irene Y. Chen, Fredrik D. Johansson, and David A. Sontag · 2018
Cited alongside, same era.
Cleaning crowdsourced labels using oracles for statistical classification
Mohamad Dolatshah, Mathew Teoh, Jiannan Wang, and Jian Pei · 2018
Cited alongside, same era.
Aurum: A data discovery system
Raul Castro Fernandez, Ziawasch Abedjan, Famien Koko, Gina Yuan, Samuel Madden, and Michael Stonebraker · 2018
Cited alongside, same era.
Introducing tensorflow hub: A library for reusable machine learning modules in tensorflow., 2018
J. Gordon · 2018
Cited alongside, same era.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor W. Tsang, and Masashi Sugiyama · 2018
Cited alongside, same era.
Data prep still dominates data scientists’ time, survey finds · 2020
Later among the works it cites.
Remixmatch: Semi-supervised learning with distribution matching and augmentation anchoring
David Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel · 2020
Later among the works it cites.
Systemds: A declarative machine learning system for the end-to-end data science lifecycle
Matthias Boehm, Iulian Antonov, Sebastian Baunsgaard, Mark Dokter, Robert Ginthör, Kevin Innerebner, Florijan Klezin, Stefanie N. Lindstaedt, Arnab Phani, Benjamin Rath, Berthold Reinwald, Shafaq Siddiqui, and Sebastian Benjamin Wrede · 2020
Later among the works it cites.
Developments in mlflow: A system to accelerate the machine learning lifecycle
Andrew Chen, Andy Chow, Aaron Davidson, Arjun DCunha, Ali Ghodsi, Sue Ann Hong, Andy Konwinski, Clemens Mewald, Siddharth Murching, Tomas Nykodym, Paul Ogilvie, Mani Parkhe, Avesh Singh, Fen Xie, Matei Zaharia, Richard Zang, Juntai Zheng, and Corey Zumar · 2020
Later among the works it cites.
Fair generative modeling via weak supervision
Kristy Choi, Aditya Grover, Trisha Singh, Rui Shu, and Stefano Ermon · 2020
Later among the works it cites.
A snapshot of the frontiers of fairness in machine learning
Alexandra Chouldechova and Aaron Roth · 2020
Later among the works it cites.
Augmix: A simple data processing method to improve robustness and uncertainty
Dan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan · 2020
Later among the works it cites.
Identifying and correcting label bias in machine learning
Heinrich Jiang and Ofir Nachum · 2020
Later among the works it cites.
Nearest neighbor classifiers over incomplete information: From certain answers to certain predictions
Bojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel, Xu Chu, Wentao Wu, and Ce Zhang · 2020
Later among the works it cites.
Fairness without demographics through adversarially reweighted learning
Preethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee, Flavien Prost, Nithum Thain, Xuezhi Wang, and Ed Chi · 2020
Later among the works it cites.
Dividemix: Learning with noisy labels as semi-supervised learning
Junnan Li, Richard Socher, and Steven C. H. Hoi · 2020
Later among the works it cites.
Secure and robust machine learning for healthcare: A survey
Adnan Qayyum, Junaid Qadir, Muhammad Bilal, and Ala Al-Fuqaha · 2020
Later among the works it cites.
Snorkel: rapid training data creation with weak supervision
Alexander Ratner, Stephen H. Bach, Henry R. Ehrenberg, Jason A. Fries, Sen Wu, and Christopher Ré · 2020
Later among the works it cites.
FR-Train: A mutual information-based approach to fair and robust training
Yuji Roh, Kangwook Lee, Steven Euijong Whang, and Changho Suh · 2020
Later among the works it cites.
Poisoning attacks on algorithmic fairness
David Solans, Battista Biggio, and Carlos Castillo · 2020
Later among the works it cites.
Learning from noisy labels with deep neural networks: A survey
Hwanjun Song, Minseok Kim, Dongmin Park, and Jae-Gil Lee · 2020
Later among the works it cites.
Robust optimization for fairness with noisy protected groups
Serena Wang, Wenshuo Guo, Harikrishna Narasimhan, Andrew Cotter, Maya R. Gupta, and Michael I. Jordan · 2020
Later among the works it cites.
Data collection and quality challenges for deep learning
Steven Euijong Whang and Jae-Gil Lee · 2020
Later among the works it cites.
Finding related tables in data lakes for interactive data science
Yi Zhang and Zachary G. Ives · 2020
Later among the works it cites.
Automated data validation in machine learning systems
Felix Biessmann, Jacek Golebiowski, Tammo Rukat, Dustin Lange, and Philipp Schmidt · 2021
Closest in time.
Validating data and models in continuous ML pipelines
Mike Dreves, Gene Huang, Zhuo Peng, Neoklis Polyzotis, Evan Rosen, and Paul Suganthan G. C · 2021
Closest in time.
Model patching: Closing the subgroup performance gap with data augmentation
Karan Goel, Albert Gu, Yixuan Li, and Christopher Ré · 2021
Closest in time.
Lightweight inspection of data preprocessing in native machine learning pipelines
Stefan Grafberger, Julia Stoyanovich, and Sebastian Schelter · 2021
Closest in time.
Inspector gadget: A data programming-based labeling system for industrial images
Geon Heo, Yuji Roh, Seonghyeon Hwang, Dayun Lee, and Steven Euijong Whang · 2021
Closest in time.
Machine learning and data cleaning: Which serves the other?
Ihab F. Ilyas and Theodoros Rekatsinas · 2021
Closest in time.
Removing spurious features can hurt accuracy and affect groups disproportionately
Fereshte Khani and Percy Liang · 2021
Closest in time.
Machine learning robustness, fairness, and their convergence
Jae-Gil Lee, Yuji Roh, Hwanjun Song, and Steven Euijong Whang · 2021
Closest in time.
CleanML: A benchmark for joint data cleaning and machine learning [experiments and analysis]
Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang · 2021
Closest in time.
On robust mean estimation under coordinate-level corruption
Zifan Liu, Jong Ho Park, Theodoros Rekatsinas, and Christos Tzamos · 2021
Closest in time.
On robust mean estimation under coordinate-level corruption
Zifan Liu, Jongho Park, Theodoros Rekatsinas, and Christos Tzamos · 2021
Closest in time.
Ease.ml: A lifecycle management system for machine learning
Leonel Aguilar Melgar, David Dao, Shaoduo Gan, Nezihe Merve Gürel, Nora Hollenstein, Jiawei Jiang, Bojan Karlas, Thomas Lemmin, Tian Li, Yang Li, Xi Rao, Johannes Rausch, Cédric Renggli, Luka Rimanic, Maurice Weber, Shuai Zhang, Zhikuan Zhao, Kevin Schawinski, Wentao Wu, and Ce Zhang · 2021
Closest in time.
From cleaning before ML to cleaning for ML
Felix Neutatz, Binger Chen, Ziawasch Abedjan, and Eugene Wu · 2021
Closest in time.
Automating data quality validation for dynamic data ingestion
Sergey Redyuk, Zoi Kaoudi, Volker Markl, and Sebastian Schelter · 2021
Closest in time.
A data quality-driven view of mlops
Cédric Renggli, Luka Rimanic, Nezihe Merve Gürel, Bojan Karlas, Wentao Wu, and Ce Zhang · 2021
Closest in time.
Fairbatch: Batch selection for model fairness
Yuji Roh, Kangwook Lee, Steven Euijong Whang, and Changho Suh · 2021
Closest in time.
Sample selection for fair and robust training
Yuji Roh, Kangwook Lee, Steven Euijong Whang, and Changho Suh · 2021
Closest in time.
JENGA - A framework to study the impact of data errors on the predictions of machine learning models
Sebastian Schelter, Tammo Rukat, and Felix Biessmann · 2021
Closest in time.
Towards out-of-distribution generalization: A survey
Zheyan Shen, Jiashuo Liu, Yue He, Xingxuan Zhang, Renzhe Xu, Han Yu, and Peng Cui · 2021
Closest in time.
Robust learning by self-transition for handling noisy labels
Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee · 2021
Closest in time.
Slice tuner: A selective data acquisition framework for accurate and fair machine learning models
Ki Hyun Tae and Steven Euijong Whang · 2021
Closest in time.
Fair classification with group-dependent label noise
Jialu Wang, Yang Liu, and Caleb Levy · 2021
Closest in time.
Enhancing the interactivity of dataframe queries by leveraging think time
Doris Xin, Devin Petersohn, Dixin Tang, Yifan Wu, Joseph E. Gonzalez, Joseph M. Hellerstein, Anthony D. Joseph, and Aditya G. Parameswaran · 2021
Closest in time.
To be robust or to be fair: Towards fairness in adversarial training
Han Xu, Xiaorui Liu, Yaxin Li, Anil K. Jain, and Jiliang Tang · 2021
Closest in time.
Omnifair: A declarative system for model-agnostic group fairness in machine learning
Hantian Zhang, Xu Chu, Abolfazl Asudeh, and Shamkant B. Navathe · 2021
Closest in time.
https://www.mturk.com/
Amazon Mechanical Turk · 2022
Closest in time.
https://aws.amazon.com/sagemaker/groundtruth/
Amazon SageMaker Ground Truth · 2022
Closest in time.
https://www.reuters.com/article/us-amazon-com-jobs-automation-insight-idUSKCN1MK08G
Amazon scraps secret AI recruiting tool that showed bias against women · 2022
Closest in time.
https://pair-code.github.io/facets/
Facets – visualization for ML datasets · 2022
Closest in time.
https://cloud.google.com/ai-platform/data-labeling/docs
GCP AI platform data labeling service · 2022
Closest in time.
https://www.bbc.com/news/technology-33347866
Google apologises for Photos app’s racist blunder · 2022
Closest in time.
https://research.samsung.com/artificial-intelligence
Principles for AI ethics · 2022
Closest in time.
https://ai.google/responsibilities/responsible-ai-practices
Responsible AI practices · 2022
Closest in time.
https://www.microsoft.com/en-us/ai/responsible-ai
Responsible AI principles from Microsoft · 2022
Closest in time.
https://www.theguardian.com/world/2021/jan/14/time-to-properly-socialise-hate-speech-ai-chatbot-pulled-from-facebook
South Korean AI chatbot pulled from Facebook after hate speech towards minorities · 2022
Closest in time.
https://www.research.ibm.com/artificial-intelligence/trusted-ai/
Trusting AI · 2022
Closest in time.
https://www.seagate.com/our-story/data-age-2025/
Data age 2025 · 2025
Closest in time.