Fetching the paper…
Reading the bibliography…
While there have been a number of remarkable breakthroughs in machine learning (ML), much of the focus has been placed on model development.
Missing data, imputation, and the bootstrap
Bradley Efron · 1994
Earlier work this paper cites.
A model of inductive bias learning
Jonathan Baxter · 2000
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence · 2008
Earlier work this paper cites.
Comparison of evaluation metrics in classification applications with imbalanced datasets
Mehrdad Fatourechi, Rabab K Ward, Steven G Mason, Jane Huggins, Alois Schlögl, and Gary E Birch · 2008
Earlier work this paper cites.
A tutorial on conformal prediction
Glenn Shafer and Vladimir Vovk · 2008
Earlier work this paper cites.
Isolation forest
Fei Tony Liu, Kai Ming Ting, and Zhi-Hua Zhou · 2008
Earlier work this paper cites.
Cluster-grouping: from subgroup discovery to clustering
Albrecht Zimmermann and Luc De Raedt · 2009
Earlier work this paper cites.
Multiple imputation by chained equations (mice): implementation in stata
Patrick Royston and Ian R White · 2011
Earlier work this paper cites.
An overview on subgroup discovery: foundations and applications
Franciso Herrera, Cristóbal José Carmona, Pedro González, and María José Del Jesus · 2011
Earlier work this paper cites.
Missforest—non-parametric missing value imputation for mixed-type data
Daniel J Stekhoven and Peter Bühlmann · 2012
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola · 2012
Earlier work this paper cites.
Hyperopt: A python library for optimizing the hyperparameters of machine learning algorithms
James Bergstra, Dan Yamins, David D Cox, et al · 2013
Earlier work this paper cites.
An instance level analysis of data complexity
Michael R Smith, Tony Martinez, and Christophe Giraud-Carrier · 2014
Earlier work this paper cites.
An instance level analysis of data complexity
Michael R Smith, Tony Martinez, and Christophe Giraud-Carrier · 2014
Earlier work this paper cites.
Training convolutional networks with noisy labels
Sainbayar Sukhbaatar, Joan Bruna, Manohar Paluri, Lubomir Bourdev, and Rob Fergus · 2014
Earlier work this paper cites.
Anomaly detection using autoencoders with nonlinear dimensionality reduction
Mayu Sakurada and Takehisa Yairi · 2014
Earlier work this paper cites.
Data wrangling: Making data useful again
Florian Endel and Harald Piringer · 2015
Earlier work this paper cites.
A review of feature selection methods with applications
Alan Jović, Karla Brkić, and Nikola Bogunović · 2015
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Yaroslav Ganin and Victor Lempitsky · 2015
Earlier work this paper cites.
Learning transferable features with deep adaptation networks
Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan · 2015
Earlier work this paper cites.
A review on evaluation metrics for data classification evaluations
Mohammad Hossin and Md Nasir Sulaiman · 2015
Earlier work this paper cites.
Variational autoencoder based anomaly detection using reconstruction probability
Jinwon An and Sungzoon Cho · 2015
Earlier work this paper cites.
Activeclean: Interactive data cleaning for statistical modeling
Sanjay Krishnan, Jiannan Wang, Eugene Wu, Michael J Franklin, and Ken Goldberg · 2016
Earlier work this paper cites.
Data programming: Creating large training sets, quickly
Alexander J Ratner, Christopher M De Sa, Sen Wu, Daniel Selsam, and Christopher Ré · 2016
Earlier work this paper cites.
Data cleaning: Overview and emerging challenges
Xu Chu, Ihab F Ilyas, Sanjay Krishnan, and Jiannan Wang · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
Loss factorization, weakly supervised learning and label noise robustness
Giorgio Patrini, Frank Nielsen, Richard Nock, and Marcello Carioni · 2016
Earlier work this paper cites.
Interpretable distribution features with maximum testing power
Wittawat Jitkrittum, Zoltán Szabó, Kacper P Chwialkowski, and Arthur Gretton · 2016
Earlier work this paper cites.
Azureml: Anatomy of a machine learning service
AzureML Team · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs
Varun Gulshan, Lily Peng, Marc Coram, Martin C Stumpe, Derek Wu, Arunachalam Narayanaswamy, Subhashini Venugopalan, Kasumi Widner, Tom Madams, Jorge Cuadros, et al · 2016
Earlier work this paper cites.
Domain adaptation for visual applications: A comprehensive survey
Gabriela Csurka · 2017
Earlier work this paper cites.
Feature selection: A data perspective
Jundong Li, Kewei Cheng, Suhang Wang, Fred Morstatter, Robert P Trevino, Jiliang Tang, and Huan Liu · 2017
Earlier work this paper cites.
Data augmentation generative adversarial networks
Antreas Antoniou, Amos Storkey, and Harrison Edwards · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Correlation alignment for unsupervised domain adaptation
Baochen Sun, Jiashi Feng, and Kate Saenko · 2017
Earlier work this paper cites.
Statistics for machine learning
Pratap Dangeti · 2017
Earlier work this paper cites.
Revisiting classifier two-sample tests
David Lopez-Paz and Maxime Oquab · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Domain invariant and class discriminative feature learning for visual domain adaptation
Shuang Li, Shiji Song, Gao Huang, Zhengming Ding, and Cheng Wu · 2018
Earlier work this paper cites.
Identifying medical diagnoses and treatable diseases by image-based deep learning
Daniel S Kermany, Michael Goldbaum, Wenjia Cai, Carolina CS Valentim, Huiying Liang, Sally L Baxter, Alex McKeown, Ge Yang, Xiaokang Wu, Fangbing Yan, et al · 2018
Earlier work this paper cites.
An introduction to domain adaptation and transfer learning
Wouter M Kouw and Marco Loog · 2018
Earlier work this paper cites.
The measure and mismeasure of fairness: A critical review of fair machine learning
Sam Corbett-Davies and Sharad Goel · 2018
Earlier work this paper cites.
Moleculenet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay Pande · 2018
Earlier work this paper cites.
Protein family-specific models using deep neural networks and transfer learning improve virtual screening and highlight the need for more data
Fergus Imrie, Anthony R Bradley, Mihaela van der Schaar, and Charlotte M Deane · 2018
Earlier work this paper cites.
Gain: Missing data imputation using generative adversarial nets
Jinsung Yoon, James Jordon, and Mihaela Schaar · 2018
Earlier work this paper cites.
Pate-gan: Generating synthetic data with differential privacy guarantees
James Jordon, Jinsung Yoon, and Mihaela Van Der Schaar · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Sherpa: Hyperparameter optimization for machine learning models
Lars Hertel, Julian Collado, Peter Sadowski, and Pierre Baldi · 2018
Cited alongside, same era.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhilu Zhang and Mert Sabuncu · 2018
Cited alongside, same era.
Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei · 2018
Cited alongside, same era.
Identifying mislabeled data using the area under the margin ranking
Geoff Pleiss, Tianyi Zhang, Ethan Elenberg, and Kilian Q Weinberger · 2020
Later among the works it cites.
Data valuation using reinforcement learning
Jinsung Yoon, Sercan Arik, and Tomas Pfister · 2020
Later among the works it cites.
Practical synthetic data generation
Khaled El Emam, Lucy Mosquera, and Richard Hoptroff · 2020
Later among the works it cites.
A comprehensive survey on transfer learning
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He · 2020
Later among the works it cites.
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré · 2020
Later among the works it cites.
Identifying mislabeled data using the area under the margin ranking
Geoff Pleiss, Tianyi Zhang, Ethan Elenberg, and Kilian Q Weinberger · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Co-teaching: Robust training of deep neural networks with extremely noisy labels
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama · 2018
Cited alongside, same era.
Iterative learning with open-set noisy labels
Yisen Wang, Weiyang Liu, Xingjun Ma, James Bailey, Hongyuan Zha, Le Song, and Shu-Tao Xia · 2018
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2018
Cited alongside, same era.
Prediction of incident hypertension within the next year: prospective study using statewide electronic health records and machine learning
Chengyin Ye, Tianyun Fu, Shiying Hao, Yan Zhang, Oliver Wang, Bo Jin, Minjie Xia, Modi Liu, Xin Zhou, Qian Wu, et al · 2018
Cited alongside, same era.
A dynamic pipeline for spatio-temporal fire risk prediction
Bhavkaran Singh Walia, Qianyi Hu, Jeffrey Chen, Fangyan Chen, Jessica Lee, Nathan Kuo, Palak Narang, Jason Batts, Geoffrey Arnold, and Michael Madaio · 2018
Cited alongside, same era.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin · 2018
Cited alongside, same era.
Later among the works it cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah Smith, and Yejin Choi · 2020
Later among the works it cites.
Data valuation using reinforcement learning
Jinsung Yoon, Sercan Arik, and Tomas Pfister · 2020
Later among the works it cites.
Anonymization through data synthesis using generative adversarial networks (ads-gan)
Jinsung Yoon, Lydia N Drumright, and Mihaela Van Der Schaar · 2020
Later among the works it cites.
Normalized loss functions for deep learning with noisy labels
Xingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano, Sarah Erfani, and James Bailey · 2020
Later among the works it cites.
Evaluating time series forecasting models: An empirical study on performance estimation methods
Vitor Cerqueira, Luis Torgo, and Igor Mozetič · 2020
Later among the works it cites.
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec · 2020
Later among the works it cites.
No subclass left behind: Fine-grained robustness in coarse-grained classification problems
Nimit Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu, and Christopher Ré · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of nlp models with checklist
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Later among the works it cites.
Incorporating machine learning and social determinants of health indicators into prospective risk adjustment for health plan payments
Jeremy A Irvin, Andrew A Kondrich, Michael Ko, Pranav Rajpurkar, Behzad Haghgoo, Bruce E Landon, Robert L Phillips, Stephen Petterson, Andrew Y Ng, and Sanjay Basu · 2020
Later among the works it cites.
Learning deep kernels for non-parametric two-sample tests
Feng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang, Arthur Gretton, and Danica J Sutherland · 2020
Later among the works it cites.
Amazon’s machine learning toolkit: Sagemaker
Ameet V Joshi · 2020
Later among the works it cites.
A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy
Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardoulakis · 2020
Later among the works it cites.
Deep learning-enabled medical computer vision
Andre Esteva, Katherine Chou, Serena Yeung, Nikhil Naik, Ali Madani, Ali Mottaghi, Yun Liu, Eric Topol, Jeff Dean, and Richard Socher · 2021
Later among the works it cites.
Machine learning to guide the use of adjuvant therapies for breast cancer
Ahmed M Alaa, Deepti Gurdasani, Adrian L Harris, Jem Rashbass, and Mihaela van der Schaar · 2021
Later among the works it cites.
A benchmark of machine learning approaches for credit score prediction
Vincenzo Moscato, Antonio Picariello, and Giancarlo Sperlí · 2021
Later among the works it cites.
Sharing learnings about our image cropping algorithm, May 2021
Rumman Chowdhury · 2021
Later among the works it cites.
Who needs mlops: What data scientists seek to accomplish and how can mlops help?
Sasu Mäkinen, Henrik Skogström, Eero Laaksonen, and Tommi Mikkonen · 2021
Later among the works it cites.
Towards mlops: A framework and maturity model
Meenu Mary John, Helena Holmström Olsson, and Jan Bosch · 2021
Later among the works it cites.
Confident learning: Estimating uncertainty in dataset labels
Curtis Northcutt, Lu Jiang, and Isaac Chuang · 2021
Later among the works it cites.
Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods
Eyke Hüllermeier and Willem Waegeman · 2021
Later among the works it cites.
Mlops: from model-centric to data-centric ai
Andrew Ng · 2021
Later among the works it cites.
What can data-centric ai learn from data and ml engineering?
Neoklis Polyzotis and Matei Zaharia · 2021
Later among the works it cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Later among the works it cites.
A data quality-driven view of mlops
Cedric Renggli, Luka Rimanic, Nezihe Merve Gürel, Bojan Karlas, Wentao Wu, and Ce Zhang · 2021
Later among the works it cites.
"everyone wants to do the model work, not the data work": Data cascades in high-stakes ai
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Kumar Paritosh, and Lora Mois Aroyo · 2021
Later among the works it cites.
Cleanml: a study for evaluating the impact of data cleaning on ml classification tasks
Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G Northcutt, Anish Athalye, and Jonas Mueller · 2021
Later among the works it cites.
Common pitfalls and recommendations for using machine learning to detect and prognosticate for covid-19 using chest radiographs and ct scans
Michael Roberts, Derek Driggs, Matthew Thorpe, Julian Gilbey, Michael Yeung, Stephan Ursprung, Angelica I Aviles-Rivero, Christian Etmann, Cathal McCague, Lucian Beer, et al · 2021
Later among the works it cites.
Synthetic data for deep learning
Sergey I Nikolenko et al · 2021
Later among the works it cites.
Sample selection for fair and robust training
Yuji Roh, Kangwook Lee, Steven Whang, and Changho Suh · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G Northcutt, Anish Athalye, and Jonas Mueller · 2021
Later among the works it cites.
Decaf: Generating fair synthetic data using causally-aware generative networks
Boris van Breugel, Trent Kyono, Jeroen Berrevoets, and Mihaela van der Schaar · 2021
Later among the works it cites.
Detecting spurious correlations with sanity tests for artificial intelligence guided radiology systems
Usman Mahmood, Robik Shrestha, David DB Bates, Lorenzo Mannelli, Giuseppe Corrias, Yusuf Emre Erdi, and Christopher Kanan · 2021
Later among the works it cites.
Counterfactual invariance to spurious correlations: Why and how to pass stress tests
Victor Veitch, Alexander D’Amour, Steve Yadlowsky, and Jacob Eisenstein · 2021
Later among the works it cites.
David Nigenda, Zohar Karnin, Muhammad Bilal Zafar, Raghu Ramesha, Alan Tan, Michele Donini, and Krishnaram Kenthapadi · 2021
Later among the works it cites.
Tabular data: Deep learning is not all you need
Ravid Shwartz-Ziv and Amitai Armon · 2022
Closest in time.
Mlops-definitions, tools and challenges
Georgios Symeonidis, Evangelos Nerantzis, Apostolos Kazakis, and George A Papakostas · 2022
Closest in time.
Temporal quality degradation in ai models
Daniel Vela, Andrew Sharp, Richard Zhang, Trang Nguyen, An Hoang, and Oleg S Pianykh · 2022
Closest in time.
From concept drift to model degradation: An overview on performance-aware drift detectors
Firas Bayram, Bestoun S Ahmed, and Andreas Kassler · 2022
Closest in time.
Advances, challenges and opportunities in creating data for trustworthy ai
Weixin Liang, Girmaw Abebe Tadesse, Daniel Ho, L Fei-Fei, Matei Zaharia, Ce Zhang, and James Zou · 2022
Closest in time.
Data-SUITE: Data-centric identification of in-distribution incongruous examples
Nabeel Seedat, Jonathan Crabbé, and Mihaela van der Schaar · 2022
Closest in time.
Conditional generation of medical time series for extrapolation to underrepresented populations
Simon Bing, Andrea Dittadi, Stefan Bauer, and Patrick Schwab · 2022
Closest in time.