Fetching the paper…
Reading the bibliography…
In recent years, machine learning models have achieved success based on the independently and identically distributed assumption.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A. (2019) · 1906
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2019) · 1907
Earlier work this paper cites.
Bootstrapping regression models
Freedman, D. A. (1981) · 1981
Earlier work this paper cites.
Neural network ensembles
Hansen, L. K. and Salamon, P. (1990) · 1990
Earlier work this paper cites.
Neural network ensembles, cross validation, and active learning
Krogh, A. and Vedelsby, J. (1994) · 1994
Earlier work this paper cites.
When networks disagree: Ensemble methods for hybrid neural networks
Perrone, M. P. and Cooper, L. N. (1995) · 1995
Earlier work this paper cites.
Bagging predictors
Breiman, L. (1996) · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and Schapire, R. E. (1997) · 1997
Earlier work this paper cites.
Popular ensemble methods: An empirical study
Opitz, D. and Maclin, R. (1999) · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
Dietterich, T. G. (2000) · 2000
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Friedman, J. H. (2001) · 2001
Earlier work this paper cites.
Ensembling neural networks: many could be better than all
Zhou, Z.-H., Wu, J., and Tang, W. (2002) · 2002
Earlier work this paper cites.
Managing diversity in regression ensembles
Brown, G., Wyatt, J. L., Tino, P., and Bengio, Y. (2005) · 2005
Earlier work this paper cites.
Boosting with early stopping: Convergence and consistency
Zhang, T. and Yu, B. (2005) · 2005
Earlier work this paper cites.
Analysis of representations for domain adaptation
Ben-David, S., Blitzer, J., Crammer, K., and Pereira, F. (2006) · 2006
Earlier work this paper cites.
Ensemble based systems in decision making
Polikar, R. (2006) · 2006
Earlier work this paper cites.
Rotation forest: A new classifier ensemble method
Rodriguez, J. J., Kuncheva, L. I., and Alonso, C. J. (2006) · 2006
Earlier work this paper cites.
In search of lost domain generalization
Gulrajani, I. and Lopez-Paz, D. (2020) · 2007
Earlier work this paper cites.
Dynamic weighted majority: An ensemble method for drifting concepts
Kolter, J. Z. and Maloof, M. A. (2007) · 2007
Earlier work this paper cites.
Covariate shift adaptation by importance weighted cross validation
Sugiyama, M., Krauledat, M., and Müller, K.-R. (2007) · 2007
Earlier work this paper cites.
Methodologies in spectral analysis of large dimensional random matrices, a review
Bai, Z. D. (2008) · 2008
Earlier work this paper cites.
Covariate shift by kernel mean matching
Gretton, A., Smola, A., Huang, J., Schmittfull, M., Borgwardt, K., and Schölkopf, B. (2008) · 2008
Earlier work this paper cites.
Discriminative learning under covariate shift
Bickel, S., Brückner, M., and Scheffer, T. (2009) · 2009
Earlier work this paper cites.
Causality
Pearl, J. (2009) · 2009
Earlier work this paper cites.
The spectrum of kernel random matrices
El Karoui, N. (2010) · 2010
Earlier work this paper cites.
Ensemble-based classifiers
Rokach, L. (2010) · 2010
Earlier work this paper cites.
A review on ensembles for the class imbalance problem: bagging-, boosting-, and hybrid-based approaches
Galar, M., Fernandez, A., Barrenechea, E., Bustince, H., and Herrera, F. (2011) · 2011
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z. and Li, Y. (2020) · 2012
Earlier work this paper cites.
Kullback-leibler divergence constrained distributionally robust optimization
Hu, Z. and Hong, L. J. (2013) · 2013
Earlier work this paper cites.
Margins, shrinkage, and boosting
Telgarsky, M. (2013) · 2013
Earlier work this paper cites.
Combining pattern classifiers: methods and algorithms
Kuncheva, L. I. (2014) · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N. (2014) · 2014
Earlier work this paper cites.
Esfahani, P. M. and Kuhn, D. (2015) · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R. R., and Srebro, N. (2015) · 2015
Cited alongside, same era.
Domain-adversarial training of neural networks
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., March, M., and Lempitsky, V. (2016) · 2016
Cited alongside, same era.
Domain adaptation with conditional transferable components
Gong, M., Zhang, K., Liu, T., Tao, D., Glymour, C., and Schölkopf, B. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Large-scale methods for distributionally robust optimization
Levy, D., Carmon, Y., Duchi, J. C., and Sidford, A. (2020) · 2020
Later among the works it cites.
Just interpolate: Kernel “ridgeless” regression can generalize
Liang, T. and Rakhlin, A. (2020) · 2020
Later among the works it cites.
Harmless interpolation of noisy data in regression
Muthukumar, V., Vodrahalli, K., Subramanian, V., and Sahai, A. (2020) · 2020
Later among the works it cites.
An investigation of why overparameterization exacerbates spurious correlations
Sagawa, S., Raghunathan, A., Koh, P. W., and Liang, P. (2020) · 2020
Later among the works it cites.
A survey of autonomous driving: Common practices and emerging technologies
Yurtsever, E., Lambert, J., Carballo, A., and Takeda, K. (2020) · 2020
Later among the works it cites.
Generalization bounds for (wasserstein) robust optimization
An, Y. and Gao, R. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2016) · 2016
Cited alongside, same era.
Stochastic gradient methods for distributionally robust optimization with f-divergences
Namkoong, H. and Duchi, J. C. (2016) · 2016
Cited alongside, same era.
Causal inference by using invariant prediction: identification and confidence intervals
Peters, J., Bühlmann, P., and Meinshausen, N. (2016) · 2016
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N. (2017) · 2017
Cited alongside, same era.
Conditional variance penalties and domain shift robustness
Heinze-Deml, C. and Meinshausen, N. (2017) · 2017
Cited alongside, same era.
Concentration inequalities and moment bounds for sample covariance operators
Koltchinskii, V. and Lounici, K. (2017) · 2017
Cited alongside, same era.
Decomposition algorithm for distributionally robust optimization using wasserstein metric
Luo, F. and Mehrotra, S. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Swad: Domain generalization by seeking flat minima
Cha, J., Chun, S., Lee, K., Cho, H.-C., Park, S., Lee, Y., and Park, S. (2021) · 2021
Later among the works it cites.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Chatterji, N. S. and Long, P. M. (2021) · 2021
Later among the works it cites.
Learning models with uniform performance via distributionally robust optimization
Duchi, J. C. and Namkoong, H. (2021) · 2021
Later among the works it cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Muthukumar, V., Narang, A., Subramanian, V., Belkin, M., Hsu, D., and Sahai, A. (2021) · 2021
Later among the works it cites.
An online method for a class of distributionally robust optimization with non-convex objectives
Qi, Q., Guo, Z., Xu, Y., Jin, R., and Yang, T. (2021) · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Later among the works it cites.
Benign overfitting in multiclass classification: All roads lead to interpolation
Wang, K., Muthukumar, V., and Thrampoulidis, C. (2021) · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2021) · 2021
Later among the works it cites.
Ensemble of averages: Improving model selection and boosting performance in domain generalization
Arpit, D., Wang, H., Zhou, Y., and Xiong, C. (2022) · 2022
Later among the works it cites.
Benign overfitting in two-layer convolutional neural networks
Cao, Y., Chen, Z., Belkin, M., and Gu, Q. (2022) · 2022
Later among the works it cites.
Calibrated ensembles can mitigate accuracy tradeoffs under distribution shift
Kumar, A., Ma, T., Liang, P., and Raghunathan, A. (2022) · 2022
Later among the works it cites.
Explicit tradeoffs between adversarial and natural distributional robustness
Moayeri, M., Banihashem, K., and Feizi, S. (2022) · 2022
Later among the works it cites.
Diverse weight averaging for out-of-distribution generalization
Rame, A., Kirchmeyer, M., Rahier, T., Rakotomamonjy, A., Gallinari, P., and Cord, M. (2022) · 2022
Later among the works it cites.
The implicit bias of benign overfitting
Shamir, O. (2022) · 2022
Later among the works it cites.
Malign overfitting: Interpolation can provably preclude invariance
Wald, Y., Yona, G., Shalit, U., and Carmon, Y. (2022) · 2022
Later among the works it cites.
Binary classification of gaussian mixtures: Abundance of support vectors, benign overfitting, and regularization
Wang, K. and Thrampoulidis, C. (2022) · 2022
Later among the works it cites.
The power and limitation of pretraining-finetuning for linear regression under covariate shift
Wu, J., Zou, D., Braverman, V., Gu, Q., and Kakade, S. (2022) · 2022
Later among the works it cites.
Sparse invariant risk minimization
Zhou, X., Lin, Y., Zhang, W., and Zhang, T. (2022) · 2022
Later among the works it cites.
The implicit bias of batch normalization in linear models and two-layer linear convolutional neural networks
Cao, Y., Zou, D., Li, Y., and Gu, Q. (2023) · 2023
Later among the works it cites.
Benign overfitting in adversarially robust linear classification
Chen, J., Cao, Y., and Gu, Q. (2023) · 2023
Later among the works it cites.
Spurious feature diversification improves out-of-distribution generalization
Lin, Y., Tan, L., Hao, Y., Wong, H., Dong, H., Zhang, W., Yang, Y., and Zhang, T. (2023) · 2023
Later among the works it cites.
Distributionally robust optimization with bias & variance reduced gradients
Mehta, R., Roulet, V., Pillutla, K., and Harchaoui, Z. (2023) · 2023
Later among the works it cites.
Simon, J. B., Karkada, D., Ghosh, N., and Belkin, M. (2023) · 2023
Later among the works it cites.
Trainable projected gradient method for robust fine-tuning
Tian, J., Dai, X., Ma, C.-Y., He, Z., Liu, Y.-C., and Kira, Z. (2023) · 2023
Later among the works it cites.
Benign overfitting in ridge regression
Tsigler, A. and Bartlett, P. L. (2023) · 2023
Later among the works it cites.
The surprising harmfulness of benign overfitting for adversarial robustness
Hao, Y. and Zhang, T. (2024) · 2024
Closest in time.
Stochastic approximation approaches to group distributionally robust optimization
Zhang, L., Zhao, P., Zhuang, Z.-H., Yang, T., and Zhou, Z.-H. (2024) · 2024
Closest in time.