Fetching the paper…
Reading the bibliography…
There often is a dilemma between ease of optimization and robust out-of-distribution (OoD) generalization.
Minimax theorems and their proofs
Simons, S · 1995
Earlier work this paper cites.
Robust Optimization , volume 28 of Princeton Series in Applied Mathematics
Ben-Tal, A., Ghaoui, L. E., and Nemirovski, A · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Uğur Güney, V., Dauphin, Y., and Bottou, L · 2017
Earlier work this paper cites.
From detection of individual metastases to classification of lymph node status at the patient level: the camelyon17 challenge
Bandi, P., Geessink, O., Manson, Q., Van Dijk, M., Balkenhol, M., Hermsen, M., Bejnordi, B. E., Lee, B., Paeng, K., Zhong, A., et al · 2018
Earlier work this paper cites.
Recognition in terra incognita
Beery, S., Van Horn, G., and Perona, P · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Trainable calibration measures for neural networks from kernel mean embeddings
Kumar, A., Sarawagi, S., and Jain, U · 2018
Earlier work this paper cites.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Cited alongside, same era.
Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2019
Cited alongside, same era.
Systematic generalisation with group invariant predictions
Ahmed, F., Bengio, Y., van Seijen, H., and Courville, A · 2020
Cited alongside, same era.
Invariant risk minimization
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 2020
Cited alongside, same era.
Self-challenging improves cross-domain generalization
Huang, Z., Wang, H., Xing, E. P., and Huang, D · 2020
Cited alongside, same era.
Out-of-distribution generalization with maximal invariant predictor
Koyama, M. and Yamaguchi, S · 2020
Environment inference for invariant learning
Creager, E., Jacobsen, J.-H., and Zemel, R · 2021
Later among the works it cites.
Domain-free adversarial splitting for domain generalization, 2021
Gu, X., Feng, J., Sun, J., and Xu, Z · 2021
Later among the works it cites.
Does invariant risk minimization capture invariance?
Kamath, P., Tangella, A., Sutherland, D. J., and Srebro, N · 2021
Later among the works it cites.
WILDS: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., Lee, T., David, E., Stavness, I., Guo, W., Earnshaw, B. A., Haque, I. S., Beery, S., Leskovec, J., Kundaje, A., Pierson, E., Levine, S., Finn, C., and Liang, P · 2021
Later among the works it cites.
Just train twice: Improving group robustness without training group information
Liu, E. Z., Haghgoo, B., Chen, A. S., Raghunathan, A., Koh, P. W., Sagawa, S., Liang, P., and Finn, C · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Out-of-distribution generalization via risk extrapolation (rex)
Krueger, D., Caballero, E., Jacobsen, J.-H., Zhang, A., Binas, J., Priol, R. L., and Courville, A · 2020
Cited alongside, same era.
Learning from failure: Training debiased classifier from biased classifier
Nam, J., Cha, H., Ahn, S., Lee, J., and Shin, J · 2020
Cited alongside, same era.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, S. O., Bengio, Y., Courville, A., Precup, D., and Lajoie, G · 2020
Cited alongside, same era.
Meta-learned invariant risk minimization
Bae, J.-H., Choi, I., and Lee, M · 2021
Cited alongside, same era.
Predict then interpolate: A simple algorithm to learn stable classifiers, 2021
Bao, Y., Chang, S., and Barzilay, R · 2021
Cited alongside, same era.
Rame, A., Dancette, C., and Cord, M · 2021
Later among the works it cites.
Gradient matching for domain generalization
Shi, Y., Seely, J., Torr, P. H., Siddharth, N., Hannun, A., Usunier, N., and Synnaeve, G · 2021
Later among the works it cites.
Evading the Simplicity Bias: Training a Diverse Set of Models Discovers Solutions with Superior OOD Generalization
Teney, D., Abbasnejad, E., Lucey, S., and van den Hengel, A · 2021
Later among the works it cites.
On calibration and out-of-domain generalization
Wald, Y., Feder, A., Greenfeld, D., and Shalit, U · 2021
Later among the works it cites.
Rosenfeld, E., Ravikumar, P., and Risteski, A · 2022
Closest in time.