Fetching the paper…
Reading the bibliography…
Machine learning models are often deployed in different settings than they were trained and validated on, posing a challenge to practitioners who wish to predict how well the deployed model will perform on a target distribution.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R. Thomas McCoy, Ellie Pavlick, and Tal Linzen · 1902
Earlier work this paper cites.
A generalization of sampling without replacement from a finite universe
D. G. Horvitz and D. J. Thompson · 1952
Earlier work this paper cites.
The central role of the propensity score in observational studies for causal effects
Paul Rosenbaum and Donald Rubin · 1983
Earlier work this paper cites.
Constructing a control group using multivariate matched sampling methods that incorporate the propensity score
Paul Rosenbaum and Donald Rubin · 1985
Earlier work this paper cites.
Improved importance sampling technique for efficient simulation of digital communication systems
D. Lu and K. Yao · 1988
Earlier work this paper cites.
Graphical Models
Steffen Lauritzen · 1996
Earlier work this paper cites.
Propensity score methods for bias reduction in the comparison of a treatment to a non-randomized control group
Ralph D’Agostino Jr · 1998
Earlier work this paper cites.
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2001
Earlier work this paper cites.
Estimation of causal effects using propensity score weighting: An application to data on right heart catheterization
Keisuke Hirano and Guido Imbens · 2001
Earlier work this paper cites.
Direct importance estimation for covariate shift adaptation
Masashi Sugiyama, Taiji Suzuki, Shinichi Nakajima, Hisashi Kashima, Paul von Bünau, and Motoaki Kawanabe · 2008
Earlier work this paper cites.
Twitter sentiment classification using distant supervision, 2009
Alec Go, Richa Bhayani, and Lei Huang · 2009
Earlier work this paper cites.
Covariate shift by kernel mean matching
Arthur Gretton, Alex Smola, Jiayuan Huang, Marcel Schmittfull, Karsten Borgwardt, and Bernhard Schölkopf · 2009
Earlier work this paper cites.
A least-squares approach to direct importance estimation
Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama · 2009
Earlier work this paper cites.
Probabilistic Graphical Models: Principles and Techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Dimensionality reduction for density ratio estimation in high-dimensional spaces
Masashi Sugiyama, Motoaki Kawanabe, and Pui Ling Chui · 2009
Earlier work this paper cites.
Direct density ratio estimation for large-scale covariate shift adaptation
Yuta Tsuboi, Hisashi Kashima, Shohei Hido, Steffen Bickel, and Masashi Sugiyama · 2009
Earlier work this paper cites.
Learning bounds for importance weighting
Corinna Cortes, Yishay Mansour, and Mehryar Mohri · 2010
Earlier work this paper cites.
Theoretical analysis of density ratio estimation
Takafumi Kanamori, Taiji Suzuki, and Masashi Sugiyama · 2010
Earlier work this paper cites.
High-dimensional Ising model selection using ℓ 1 \ell_{1} -regularized logistic regression
Pradeep Ravikumar, Martin Wainwright, and John Lafferty · 2010
Earlier work this paper cites.
An introduction to propensity score methods for reducing the effects of confounding in observational studies
Peter Austin · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond Daly, Peter Pham, Dan Huang, Andrew Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Active learning
Burr Settles · 2012
Earlier work this paper cites.
Robust solutions of optimization problems affected by uncertain probabilities
Aharon Ben-Tal, Dick den Hertog, Anja De Waegenaere, Bertrand Melenberg, and Gijs Rennen · 2013
Earlier work this paper cites.
Structure estimation for discrete graphical models: Generalized covariance matrices and their inverses
Po-Ling Loh and Martin Wainwright · 2013
Cited alongside, same era.
Monte Carlo theory, methods, and examples
Art Owen · 2013
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel Bowman, Gabor Angeli, Christopher Potts, and Christopher Manning · 2015
Cited alongside, same era.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Cited alongside, same era.
Balancing covariates via propensity score weighting
Fan Li, Kari Lock Morgan, and Alan Zaslavsky · 2016
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Slice finder: Automated data slicing for model validation
Neoklis Polyzotis, Steven Whang, Tim Kraska, and Yeounoh Chung · 2019
Later among the works it cites.
Low-dimensional density ratio estimation for covariate shift correction
Petar Stojanov, Mingming Gong, Jaime Carbonell, and Kun Zhang · 2019
Later among the works it cites.
Rethinking importance weighting for deep learning under distribution shift
Tongtong Fang, Nan Lu, Gang Niu, and Masashi Sugiyama · 2020
Later among the works it cites.
Towards principled unskewing: Viewing 2020 election polls through a corrective lens from 2016
Michael Isakov and Shiro Kuriwaki · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsupervised domain adaptation with residual transfer networks
Mingsheng Long, Han Zhu, Jianmin Wang, and Michael I. Jordan · 2016
Cited alongside, same era.
Learning in implicit generative models
S. Mohamed and Balaji Lakshminarayanan · 2016
Cited alongside, same era.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2017
Cited alongside, same era.
Lipschitz Density-Ratios, Structured Data, and Data-driven Tuning
Samory Kpotufe · 2017
Cited alongside, same era.
Trimmed density ratio estimation
Song Liu, Akiko Takeda, Taiji Suzuki, and Kenji Fukumizu · 2017
Cited alongside, same era.
Does distributionally robust supervised learning give robust classifiers?
Weihua Hu, Gang Niu, Issei Sato, and Masashi Sugiyama · 2018
Cited alongside, same era.
Learning the difference that makes a difference with counterfactually augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton · 2020
Later among the works it cites.
Robust importance weighting for covariate shift
Fengpei Li, Henry Lam, and Siddharth Prusty · 2020
Later among the works it cites.
Posterior ratio estimation of latent variables
Song Liu, Yulong Zhang, Mingxuan Yi, and Mladen Kolar · 2020
Later among the works it cites.
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré · 2020
Later among the works it cites.
Covariate shift adaptation in high-dimensional and divergent distributions
Felipe Maia Polo and Renato Vicente · 2020
Later among the works it cites.
Efficient test collection construction via active learning
Md Mustafizur Rahman, Mucahid Kutlu, Tamer Elsayed, and Matthew Lease · 2020
Later among the works it cites.
Telescoping density-ratio estimation
Benjamin Rhodes, Kai Xu, and Michael Gutmann · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Later among the works it cites.
Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization
Shiori Sagawa, Pang Wei Koh, Tatsunori Hashimoto, and Percy Liang · 2020
Later among the works it cites.
Not your grandfathers test set: Reducing labeling effort for testing
Begum Taskazan, Jiri Navratil, Matthew Arnold, Anupama Murthi, Ganesh Venkataraman, and Benjamin Elder · 2020
Later among the works it cites.
Statistics of robust optimization: A generalized empirical likelihood approach
John Duchi, Peter Glynn, and Hongseok Namkoong · 2021
Closest in time.
Robustness gym: Unifying the NLP evaluation landscape
Karan Goel, Nazneen Rajani, Jesse Vig, Samson Tan, Jason Wu, Stephan Zheng, Caiming Xiong, Mohit Bansal, and Christopher Ré · 2021
Closest in time.
Model performance scaling with multiple data sources
Tatsunori Hashimoto · 2021
Closest in time.
Two-sample inference for high-dimensional Markov networks
Byol Kim, Song Liu, and Mladen Kolar · 2021
Closest in time.
WILDS: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang · 2021
Closest in time.
SliceLine: Fast, linear-algebra-based slice finding for ML model debugging
Svetlana Sagadeeva and Matthias Boehm · 2021
Closest in time.
Evaluating model robustness and stability to dataset shift
Adarsh Subbaswamy, Roy Adams, and Suchi Saria · 2021
Closest in time.
Slice tuner: A selective data acquisition framework for accurate and fair machine learning models
Ki Hyun Tae and Steven Whang · 2021
Closest in time.