Fetching the paper…
Reading the bibliography…
We study why overparameterization -- increasing model size well beyond the point of zero training error -- can hurt test error on minority groups despite improving average test error when there are spurious correlations in the data.
Statistical mechanics of learning: Generalization
Opper, M · 1995
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Shimodaira, H · 2000
Earlier work this paper cites.
The class imbalance problem: A systematic study
Japkowicz, N. and Stephen, S · 2002
Earlier work this paper cites.
Margin maximizing loss functions
Rosset, S., Zhu, J., and Hastie, T. J · 2004
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S · 2011
Earlier work this paper cites.
Fairness through awareness
Dwork, C., Hardt, M., Pitassi, T., Reingold, O., and Zemel, R · 2012
Earlier work this paper cites.
Robust solutions of optimization problems affected by uncertain probabilities
Ben-Tal, A., den Hertog, D., Waegenaere, A. D., Melenberg, B., and Rennen, G · 2013
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Robust learning under uncertain test distributions: Relating covariate shift to model misspecification
Wen, J., Yu, C., and Greiner, R · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Blodgett, S. L., Green, L., and O’Connor, B · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Hardt, M., Price, E., and Srebo, N · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S. and Saxe, A. M · 2017
Cited alongside, same era.
Learning from class-imbalanced data: Review of methods and applications
Haixiang, G., Yijing, L., Shang, J., Mingyun, G., Yuanyue, H., and Bing, G · 2017
Cited alongside, same era.
Inherent trade-offs in the fair determination of risk scores
Kleinberg, J., Mullainathan, S., and Raghavan, M · 2017
Cited alongside, same era.
Variance regularization with convex objectives
Namkoong, H. and Duchi, J · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Later among the works it cites.
What is the effect of importance weighting in deep learning?
Byrd, J. and Lipton, Z · 2019
Later among the works it cites.
Learning imbalanced datasets with label-distribution-aware margin loss
Cao, K., Wei, C., Gaidon, A., Arechiga, N., and Ma, T · 2019
Later among the works it cites.
Class-balanced loss based on effective number of samples
Cui, Y., Jia, M., Lin, T., Song, Y., and Belongie, S · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2019
Later among the works it cites.
Benchmarking neural network robustness to common corruptions and perturbations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Places: A 10 million image database for scene recognition
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A · 2017
Cited alongside, same era.
A systematic study of the class imbalance problem in convolutional neural networks
Buda, M., Maki, A., and Mazurowski, M. A · 2018
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Buolamwini, J. and Gebru, T · 2018
Cited alongside, same era.
Fairness without demographics in repeated loss minimization
Hashimoto, T. B., Srivastava, M., Namkoong, H., and Liang, P · 2018
Cited alongside, same era.
Does distributionally robust supervised learning give robust classifiers?
Hu, W., Niu, G., Sato, I., and Sugiyama, M · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
Hendrycks, D. and Dietterich, T · 2019
Later among the works it cites.
Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D · 2019
Later among the works it cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, R. T., Pavlick, E., and Linzen, T · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Later among the works it cites.
Distributionally robust language modeling
Oren, Y., Sagawa, S., Hashimoto, T., and Liang, P · 2019
Later among the works it cites.
Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2020
Closest in time.
Rethinking bias-variance trade-off for generalization of neural networks
Yang, Z., Yu, Y., You, C., Steinhardt, J., and Ma, Y · 2020
Closest in time.