Fetching the paper…
Reading the bibliography…
Statistical tests for dataset shift are susceptible to false alarms: they are sensitive to minor differences when there is in fact adequate sample coverage and predictive performance.
The irises of the gaspe peninsula
Edgar Anderson · 1935
Earlier work this paper cites.
The impact of changing populations on classifier performance
Mark G Kelly, David J Hand, and Niall M Adams · 1999
Earlier work this paper cites.
Powerful goodness-of-fit tests based on the likelihood ratio
Jin Zhang · 2002
Earlier work this paper cites.
Testing for equal distributions in high dimension
Gábor J Székely, Maria L Rizzo, et al · 2004
Earlier work this paper cites.
A new evaluation measure for imbalanced datasets
Cheng G Weng and Josiah Poon · 2008
Earlier work this paper cites.
A framework for monitoring classifiers’ performance: when and why failure occurs?
David A Cieslak and Nitesh V Chawla · 2009
Earlier work this paper cites.
Auc optimization and the two-sample problem
Stéphan Clémençon, Marine Depecker, and Nicolas Vayatis · 2009
Earlier work this paper cites.
Sequential implementation of monte carlo tests with uniformly bounded resampling risk
Axel Gandy · 2009
Earlier work this paper cites.
Measuring classifier performance: a coherent alternative to the area under the roc curve
David J Hand · 2009
Earlier work this paper cites.
Loop: local outlier probabilities
Hans-Peter Kriegel, Peer Kröger, Erich Schubert, and Arthur Zimek · 2009
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence · 2009
Earlier work this paper cites.
Weighted area under the receiver operating characteristic curve and its application to gene selection
Jialiang Li and Jason P Fine · 2010
Earlier work this paper cites.
Testing statistical hypotheses of equivalence and noninferiority
Stefan Wellek · 2010
Earlier work this paper cites.
Interpreting and unifying outlier scores
Hans-Peter Kriegel, Peer Kroger, Erich Schubert, and Arthur Zimek · 2011
Earlier work this paper cites.
Study on the impact of partition-induced dataset shift on k k -fold cross-validation
Jose García Moreno-Torres, José A Sáez, and Francisco Herrera · 2012
Earlier work this paper cites.
A bayesian wilcoxon signed-rank test based on the dirichlet process
Alessio Benavoli, Giorgio Corani, Francesca Mangili, Marco Zaffalon, and Fabrizio Ruggeri · 2014
Cited alongside, same era.
Robust learning under uncertain test distributions: Relating covariate shift to model misspecification
Junfeng Wen, Chun-Nam Yu, and Russell Greiner · 2014
Cited alongside, same era.
The p-value you can’t buy
Eugene Demidenko · 2016
Cited alongside, same era.
Estimation stability with cross-validation (escv)
Chinghway Lim and Bin Yu · 2016
Cited alongside, same era.
Linking losses for density ratio and class-probability estimation
Aditya Menon and Cheng Soon Ong · 2016
Cited alongside, same era.
The asa’s statement on p-values: context, process, and purpose
Ronald L Wasserstein, Nicole A Lazar, et al · 2016
Cited alongside, same era.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Jasper Snoek, Yaniv Ovadia, Emily Fertig, Balaji Lakshminarayanan, Sebastian Nowozin, D Sculley, Joshua Dillon, Jie Ren, and Zachary Nado · 2019
Later among the works it cites.
Two-sample test based on classification probability
Haiyan Cai, Bryan Goggin, and Qingtang Jiang · 2020
Later among the works it cites.
Goodness-of-fit testing in high dimensional generalized linear models
Jana Janková, Rajen D Shah, Peter Bühlmann, and Richard J Samworth · 2020
Later among the works it cites.
Challenges in deploying machine learning: a survey of case studies
Andrei Paleyes, Raoul-Gabriel Urma, and Neil D Lawrence · 2020
Later among the works it cites.
Big data? statistical process control can help!
Peihua Qiu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openml: An r package to connect to the machine learning platform openml
Giuseppe Casalicchio, Jakob Bossek, Michel Lang, Dominik Kirchhoff, Pascal Kerschke, Benjamin Hofner, Heidi Seibold, Joaquin Vanschoren, and Bernd Bischl · 2017
Cited alongside, same era.
Superheat: An r package for creating beautiful and extendable heatmaps for visualizing complex data
Rebecca L Barter and Bin Yu · 2018
Cited alongside, same era.
To trust or not to trust a classifier
Heinrich Jiang, Been Kim, Melody Y Guan, and Maya R Gupta · 2018
Cited alongside, same era.
Detecting and correcting for label shift with black box predictors
Zachary Lipton, Yu-Xiang Wang, and Alexander Smola · 2018
Cited alongside, same era.
The p-value requires context, not a threshold
Rebecca A Betensky · 2019
Cited alongside, same era.
Valid p-values behave exactly as they should: Some misleading criticisms of p-values and their resolution with s-values
Sander Greenland · 2019
Cited alongside, same era.
Jie M Zhang, Mark Harman, Lei Ma, and Yang Liu · 2020
Later among the works it cites.
Testing for outliers with conformal p-values
Stephen Bates, Emmanuel Candès, Lihua Lei, Yaniv Romano, and Matteo Sesia · 2021
Closest in time.
Confidence-based out-of-distribution detection: a comparative study and analysis
Christoph Berger, Magdalini Paschali, Ben Glocker, and Konstantinos Kamnitsas · 2021
Closest in time.
Leave-one-out kernel density estimates for outlier detection
Sevvandi Kandanaarachchi and Rob J Hyndman · 2021
Closest in time.
Classification accuracy as a proxy for two-sample testing
Ilmun Kim, Aaditya Ramdas, Aarti Singh, and Larry Wasserman · 2021
Closest in time.
Density of states estimation for out of distribution detection
Warren Morningstar, Cusuh Ham, Andrew Gallagher, Balaji Lakshminarayanan, Alex Alemi, and Joshua Dillon · 2021
Closest in time.
Evaluating model robustness and stability to dataset shift
Adarsh Subbaswamy, Roy Adams, and Suchi Saria · 2021
Closest in time.
Retrain or not retrain: Conformal test martingales for change-point detection
Vladimir Vovk, Ivan Petej, Ilia Nouretdinov, Ernst Ahlberg, Lars Carlsson, and Alex Gammerman · 2021
Closest in time.
On the use of random forest for two-sample testing
Simon Hediger, Loris Michel, and Jeffrey Näf · 2022
Closest in time.
Tracking the risk of a deployed model and detecting harmful distribution shifts
Aleksandr Podkopaev and Aaditya Ramdas · 2022
Closest in time.