Fetching the paper…
Reading the bibliography…
Covariate-shift generalization, a typical case in out-of-distribution (OOD) generalization, requires a good performance on the unknown test distribution, which varies from the accessible training distribution in the form of covariate shift.
The central role of the propensity score in observational studies for causal effects
Rosenbaum, P. R. and Rubin, D. B · 1983
Earlier work this paper cites.
Statistics and causal inference
Holland, P. W · 1986
Earlier work this paper cites.
A practical approach to feature selection
Kira, K. and Rendell, L. A · 1992
Earlier work this paper cites.
Irrelevant features and the subset selection problem
John, G. H., Kohavi, R., and Pfleger, K · 1994
Earlier work this paper cites.
Selection of relevant features in machine learning
Langley, P. et al · 1994
Earlier work this paper cites.
On the sensitivity of solution components in linear systems of equations
Chandrasekaran, S. and Ipsen, I. C · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R · 1996
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Shimodaira, H · 2000
Earlier work this paper cites.
Causation, prediction, and search
Spirtes, P., Glymour, C. N., Scheines, R., and Heckerman, D · 2000
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Friedman, J. H · 2001
Earlier work this paper cites.
Optimal structure identification with greedy search
Chickering, D. M · 2002
Earlier work this paper cites.
An introduction to variable and feature selection
Guyon, I. and Elisseeff, A · 2003
Earlier work this paper cites.
Variable selection using svm-based criteria
Rakotomamonjy, A · 2003
Earlier work this paper cites.
Towards principled feature selection: Relevancy, filters and wrappers
Tsamardinos, I. and Aliferis, C. F · 2003
Earlier work this paper cites.
Simultaneous feature selection and clustering using mixture models
Law, M. H., Figueiredo, M. A., and Jain, A. K · 2004
Earlier work this paper cites.
Causal discovery using a bayesian local causal discovery algorithm
Mani, S. and Cooper, G. F · 2004
Earlier work this paper cites.
Causal inference using potential outcomes: Design, modeling, decisions
Rubin, D. B · 2005
Earlier work this paper cites.
Learning bounds for kernel regression using effective data dimensionality
Zhang, T · 2005
Earlier work this paper cites.
Regularization and variable selection via the elastic net
Zou, H. and Hastie, T · 2005
Earlier work this paper cites.
Pattern recognition
Bishop, C. M · 2006
Earlier work this paper cites.
Multi criteria wrapper improvements to naive bayes learning
Cortizo, J. C. and Giraldez, I · 2006
Earlier work this paper cites.
Domain adaptation for statistical classifiers
Daume III, H. and Marcu, D · 2006
Earlier work this paper cites.
Gene selection and classification of microarray data using random forest
Díaz-Uriarte, R. and De Andres, S. A · 2006
Earlier work this paper cites.
Correcting sample selection bias by unlabeled data
Huang, J., Gretton, A., Borgwardt, K., Schölkopf, B., and Smola, A · 2006
Earlier work this paper cites.
Balance-subsampled stable prediction
Kuang, K., Zhang, H., Wu, F., Zhuang, Y., and Zhang, A · 2006
Earlier work this paper cites.
Analysis of representations for domain adaptation
Ben-David, S., Blitzer, J., Crammer, K., Pereira, F., et al · 2007
Earlier work this paper cites.
Discriminative learning for differing training and test distributions
Bickel, S., Brückner, M., and Scheffer, T · 2007
Earlier work this paper cites.
Kernel measures of conditional dependence
Fukumizu, K., Gretton, A., Sun, X., and Schölkopf, B · 2007
Cited alongside, same era.
Towards scalable and data efficient learning of markov boundaries
Pena, J. M., Nilsson, R., Björkegren, J., and Tegnér, J · 2007
Cited alongside, same era.
Random features for large-scale kernel machines
Rahimi, A., Recht, B., et al · 2007
Cited alongside, same era.
Mixture regression for covariate shift
Storkey, A. J. and Sugiyama, M · 2007
Cited alongside, same era.
Direct importance estimation for covariate shift adaptation
Sugiyama, M., Suzuki, T., Nakajima, S., Kashima, H., von Bünau, P., and Kawanabe, M · 2008
Cited alongside, same era.
A least-squares approach to direct importance estimation
Kanamori, T., Hido, S., and Sugiyama, M · 2009
Cited alongside, same era.
Markov boundary discovery with ridge regularized linear models
Strobl, E. V. and Visweswaran, S · 2016
Later among the works it cites.
Generalized additive models
Hastie, T. J. and Tibshirani, R. J · 2017
Later among the works it cites.
Generalized score functions for causal discovery
Huang, B., Zhang, K., Lin, Y., Schölkopf, B., and Glymour, C · 2018
Later among the works it cites.
Stable prediction across unknown environments
Kuang, K., Cui, P., Athey, S., Xiong, R., and Li, B · 2018
Later among the works it cites.
Causally regularized learning with agnostic data selection bias
Shen, Z., Cui, P., Kuang, K., Li, B., and Chen, P · 2018
Later among the works it cites.
Relief-based feature selection: Introduction and review
Urbanowicz, R. J., Meeker, M., La Cava, W., Olson, R. S., and Moore, J. H · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A comparison of random forest and its gini importance with standard chemometric methods for the feature selection and classification of spectral data
Menze, B. H., Kelm, B. M., Masuch, R., Himmelreich, U., Bachert, P., Petrich, W., and Hamprecht, F. A · 2009
Cited alongside, same era.
Sparse additive models
Ravikumar, P., Lafferty, J., Liu, H., and Wasserman, L · 2009
Cited alongside, same era.
A density-ratio framework for statistical data processing
Sugiyama, M., Kanamori, T., Suzuki, T., Hido, S., Sese, J., Takeuchi, I., and Wang, L · 2009
Cited alongside, same era.
Classification and regression trees
Loh, W.-Y · 2011
Cited alongside, same era.
Matrix analysis
Horn, R. A. and Johnson, C. R · 2012
Cited alongside, same era.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Cited alongside, same era.
Dags with no tears: Continuous optimization for structure learning
Zheng, X., Aragam, B., Ravikumar, P. K., and Xing, E. P · 2018
Later among the works it cites.
Approximate kernel-based conditional independence tests for fast non-parametric causal discovery
Strobl, E. V., Zhang, K., and Visweswaran, S · 2019
Later among the works it cites.
Focused context balancing for robust offline policy evaluation
Zou, H., Kuang, K., Chen, B., Chen, P., and Cui, P · 2019
Later among the works it cites.
Distributionally robust losses for latent covariate mixtures
Duchi, J., Hashimoto, T., and Namkoong, H · 2020
Later among the works it cites.
Rethinking importance weighting for deep learning under distribution shift
Fang, T., Lu, N., Niu, G., and Sugiyama, M · 2020
Later among the works it cites.
A unified view of label shift estimation
Garg, S., Wu, Y., Balakrishnan, S., and Lipton, Z. C · 2020
Later among the works it cites.
The hardness of conditional independence testing and the generalised covariance measure
Shah, R. D. and Peters, J · 2020
Later among the works it cites.
Stable learning via sample reweighting
Shen, Z., Cui, P., Zhang, T., and Kuang, K · 2020
Later among the works it cites.
Feature selection using stochastic gates
Yamada, Y., Lindenbaum, O., Negahban, S., and Kluger, Y · 2020
Later among the works it cites.
Learning sparse nonparametric dags
Zheng, X., Dan, C., Aragam, B., Ravikumar, P., and Xing, E · 2020
Later among the works it cites.
Counterfactual prediction for bundle treatment
Zou, H., Cui, P., Li, B., Shen, Z., Ma, J., Yang, H., and He, Y · 2020
Later among the works it cites.
Learning models with uniform performance via distributionally robust optimization
Duchi, J. C. and Namkoong, H · 2021
Closest in time.
Daring: Differentiable causal discovery with residual independence
He, Y., Cui, P., Shen, Z., Xu, R., Liu, F., and Jiang, Y · 2021
Closest in time.
Out-of-distribution generalization via risk extrapolation (rex)
Krueger, D., Caballero, E., Jacobsen, J.-H., Zhang, A., Binas, J., Zhang, D., Le Priol, R., and Courville, A · 2021
Closest in time.
Stabilizing variable selection and regression
Pfister, N., Williams, E. G., Peters, J., Aebersold, R., and Bühlmann, P · 2021
Closest in time.
Optimal representations for covariate shift
Ruan, Y., Dubois, Y., and Maddison, C. J · 2021
Closest in time.
Towards out-of-distribution generalization: A survey
Shen, Z., Liu, J., He, Y., Zhang, X., Xu, R., Yu, H., and Cui, P · 2021
Closest in time.
Causalvae: disentangled representation learning via neural structural causal models
Yang, M., Liu, F., Chen, Z., Shen, X., Hao, J., and Wang, J · 2021
Closest in time.
Deep stable learning for out-of-distribution generalization
Zhang, X., Cui, P., Xu, R., Zhou, L., He, Y., and Shen, Z · 2021
Closest in time.
Domain generalization: A survey
Zhou, K., Liu, Z., Qiao, Y., Xiang, T., and Loy, C. C · 2021
Closest in time.
Stable learning establishes some common ground between causal inference and machine learning
Cui, P. and Athey, S · 2022
Closest in time.