Fetching the paper…
Reading the bibliography…
Learning robust models that generalize well under changes in the data distribution is critical for real-world applications.
On the mathematical foundations of theoretical statistics
Fisher, R. A · 1922
Earlier work this paper cites.
Information and accuracy attainable in the estimation of statistical parameters
C.R., R · 1945
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Becker, S. and Le Cun, Y · 1988
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J., Solla, S., Howard, R., and Jackel, L · 1990
Earlier work this paper cites.
An overview of statistical learning theory
Vapnik, V. N · 1999
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
Heskes, T · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N · 2002
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
Mnist handwritten digit database, 2010
LeCun, Y., Cortes, C., and Burges, C · 2010
Earlier work this paper cites.
Generalizing from several related classification tasks to a new unlabeled sample
Blanchard, G., Lee, G., and Scott, C · 2011
Earlier work this paper cites.
The beginning of infinity: Explanations that transform the world
Deutsch, D · 2011
Earlier work this paper cites.
Improving first and second-order methods by modeling uncertainty
Le Roux, N., Bengio, Y., and Fitzgibbon, A · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias
Fang, C., Xu, Y., and Rockmore, D. N · 2013
Earlier work this paper cites.
Domain generalization via invariant feature representation
Muandet, K., Balduzzi, D., and Schölkopf, B · 2013
Earlier work this paper cites.
No more pesky learning rates
Schaul, T., Zhang, S., and LeCun, Y · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Transfer joint matching for unsupervised domain adaptation
Long, M., Wang, J., Ding, G., Sun, J., and Yu, P. S · 2014
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
Martens, J · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
Tzeng, E., Hoffman, J., Zhang, N., Saenko, K., and Darrell, T · 2014
Earlier work this paper cites.
Domain generalization for object recognition with multi-task autoencoders
Ghifary, M., Kleijn, W. B., Zhang, M., and Balduzzi, D · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Domain-adversarial training of neural networks
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., and Lempitsky, V · 2016
Earlier work this paper cites.
Domain adaptation with conditional transferable components
Gong, M., Zhang, K., Liu, T., Tao, D., Glymour, C., and Schölkopf, B · 2016
Earlier work this paper cites.
Causal inference by using invariant prediction: identification and confidence intervals
Peters, J., Bühlmann, P., and Meinshausen, N · 2016
Earlier work this paper cites.
Deep coral: Correlation alignment for deep domain adaptation
Sun, B. and Saenko, K · 2016
Earlier work this paper cites.
Return of frustratingly easy domain adaptation
Sun, B., Feng, J., and Saenko, K · 2016
Earlier work this paper cites.
Deep domain generalization with structured low-rank constraint
Ding, Z. and Fu, Y · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Deeper, broader and artier domain generalization
Li, D., Yang, Y., Song, Y.-Z., and Hospedales, T. M · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
Lopez-Paz, D. and Ranzato, M. A · 2017
Earlier work this paper cites.
Deep hashing network for unsupervised domain adaptation
Venkateswara, H., Eusebio, J., Chakraborty, S., and Panchanathan, S · 2017
Earlier work this paper cites.
Recognition in terra incognita
Beery, S., Van Horn, G., and Perona, P · 2018
Earlier work this paper cites.
Adapting auxiliary losses using gradient similarity
Du, Y., Czarnecki, W. M., Jayakumar, S. M., Farajtabar, M., Pascanu, R., and Lakshminarayanan, B · 2018
Earlier work this paper cites.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Three factors influencing minima in SGD
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Storkey, A., and Bengio, Y · 2018
Cited alongside, same era.
Best sources forward: domain generalization through source-specific nets
Mancini, M., Bulo, S. R., Caputo, B., and Ricci, E · 2018
Cited alongside, same era.
Invariant models for causal transfer learning
Rojas-Carulla, M., Schölkopf, B., Turner, R., and Peters, J · 2018
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks, 2018
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2018
Woodfisher: Efficient second-order approximation for neural network compression
Singh, S. P. and Alistarh, D · 2020
Later among the works it cites.
Unshuffling data for improved generalization
Teney, D., Abbasnejad, E., and van den Hengel, A · 2020
Later among the works it cites.
On the interplay between noise and curvature and its effect on optimization and generalization
Thomas, V., Pedregosa, F., van Merriënboer, B., Manzagol, P.-A., Bengio, Y., and Roux, N. L · 2020
Later among the works it cites.
Heterogeneous domain generalization via domain mixup
Wang, Y., Li, H., and Kot, A. C · 2020
Later among the works it cites.
Dual mixup regularized learning for adversarial domain adaptation
Wu, Y., Inkpen, D., and El-Roby, A · 2020
Later among the works it cites.
Improve unsupervised domain adaptation with mixup training
Yan, S., Song, H., Li, N., Zou, L., and Ren, L · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Building machines that learn and think like people
Tenenbaum, J · 2018
Cited alongside, same era.
Faster gaze prediction with dense networks and fisher pruning
Theis, L., Korshunova, I., Tejani, A., and Huszár, F · 2018
Cited alongside, same era.
Gradient diversity: a key ingredient for scalable distributed learning
Yin, D., Pananjady, A., Lam, M., Papailiopoulos, D., Ramchandran, K., and Bartlett, P · 2018
Cited alongside, same era.
Invariance principle meets information bottleneck for out-of-distribution generalization
Ahuja, K., Caballero, E., Zhang, D., Bengio, Y., Mitliagkas, I., and Rish, I · 2019
Cited alongside, same era.
Invariant risk minimization
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 2019
Cited alongside, same era.
Input similarity from the neural network perspective
Charpiat, G., Girard, N., Felardos, L., and Tarabalka, Y · 2019
Cited alongside, same era.
Later among the works it cites.
Gradient surgery for multi-task learning
Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C · 2020
Later among the works it cites.
Adaptive risk minimization: A meta-learning approach for tackling group distribution shift
Zhang, M., Marklund, H., Dhawan, N., Gupta, A., Levine, S., and Finn, C · 2020
Later among the works it cites.
Systematic generalisation with group invariant predictions
Ahmed, F., Bengio, Y., van Seijen, H., and Courville, A · 2021
Closest in time.
Domain generalization by marginal transfer learning
Blanchard, G., Deshmukh, A. A., Dogan, U., Lee, G., and Scott, C · 2021
Closest in time.
SWAD: Domain generalization by seeking flat minima
Cha, J., Chun, S., Lee, K., Cho, H.-C., Park, S., Lee, Y., and Park, S · 2021
Closest in time.
Environment inference for invariant learning
Creager, E., Jacobsen, J.-H., and Zemel, R · 2021
Closest in time.
Vivit: Curvature access through the generalized gauss-newton’s low-rank structure
Dangel, F., Tatzel, L., and Hennig, P · 2021
Closest in time.
Ai for radiographic covid-19 detection selects shortcuts over signal
DeGrave, A. J., Janizek, J. D., and Lee, S.-I · 2021
Closest in time.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Closest in time.
Efficient matrix-free approximations of second-order information, with applications to pruning and optimization
Frantar, E., Kurtic, E., and Alistarh, D · 2021
Closest in time.
In search of lost domain generalization
Gulrajani, I. and Lopez-Paz, D · 2021
Closest in time.
Out-of-distribution prediction with invariant risk minimization: The limitation and an effective fix
Guo, R., Zhang, P., Liu, H., and Kiciman, E · 2021
Closest in time.
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
Jastrzebski, S., Arpit, D., Astrand, O., Kerg, G. B., Wang, H., Xiong, C., Socher, R., Cho, K., and Geras, K. J · 2021
Closest in time.
Does invariant risk minimization capture invariance?
Kamath, P., Tangella, A., Sutherland, D., and Srebro, N · 2021
Closest in time.
Out-of-distribution generalization via risk extrapolation (rex)
Krueger, D., Caballero, E., Jacobsen, J.-H., Zhang, A., Binas, J., Zhang, D., Priol, R. L., and Courville, A · 2021
Closest in time.
Group fisher pruning for practical network compression
Liu, L., Zhang, S., Kuang, Z., Zhou, A., Xue, J.-H., Wang, X., Chen, Y., Yang, W., Liao, Q., and Zhang, W · 2021
Closest in time.
Domain generalization via gradient surgery
Mansilla, L., Echeveste, R., Milone, D. H., and Ferrante, E · 2021
Closest in time.
Reducing domain gap by reducing style bias
Nam, H., Lee, H., Park, J., Yoon, W., and Yoo, D · 2021
Closest in time.
Learning explanations that are hard to vary
Parascandolo, G., Neitz, A., Orvieto, A., Gresele, L., and Schölkopf, B · 2021
Closest in time.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, S.-O., Bengio, Y., Courville, A., Precup, D., and Lajoie, G · 2021
Closest in time.
Common pitfalls and recommendations for using machine learning to detect and prognosticate for covid-19 using chest radiographs and ct scans
Roberts, M., Driggs, D., Thorpe, M., Gilbey, J., Yeung, M., Ursprung, S., Aviles-Rivero, A. I., Etmann, C., McCague, C., Beer, L., et al · 2021
Closest in time.
The risks of invariant risk minimization
Rosenfeld, E., Ravikumar, P. K., and Risteski, A · 2021
Closest in time.
Sand-mask: An enhanced gradient masking strategy for the discovery of invariances in domain generalization
Shahtalebi, S., Gagnon-Audet, J.-C., Laleh, T., Faramarzi, M., Ahuja, K., and Rish, I · 2021
Closest in time.
Gradient matching for domain generalization
Shi, Y., Seely, J., Torr, P. H., Siddharth, N., Hannun, A., Usunier, N., and Synnaeve, G · 2021
Closest in time.
Evading the simplicity bias: Training a diverse set of models discovers solutions with superior OOD generalization
Teney, D., Abbasnejad, E., Lucey, S., and van den Hengel, A · 2021
Closest in time.
Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020
Turner, R., Eriksson, D., McCourt, M., Kiili, J., Laaksonen, E., Xu, Z., and Guyon, I · 2021
Closest in time.
On calibration and out-of-domain generalization
Wald, Y., Feder, A., Greenfeld, D., and Shalit, U · 2021
Closest in time.
In-n-out: Pre-training and self-training using auxiliary information for out-of-distribution robustness
Xie, S. M., Kumar, A., Jones, R., Khani, F., Ma, T., and Liang, P · 2021
Closest in time.
Ood-bench: Benchmarking and understanding out-of-distribution generalization datasets and algorithms
Ye, N., Li, K., Hong, L., Bai, H., Chen, Y., Zhou, F., and Li, Z · 2021
Closest in time.
Deep stable learning for out-of-distribution generalization
Zhang, X., Cui, P., Xu, R., Zhou, L., He, Y., and Shen, Z · 2021
Closest in time.
Diverse weight averaging for out-of-distribution generalization
Rame, A., Kirchmeyer, M., Rahier, T., Rakotomamonjy, A., Gallinari, P., and Cord, M · 2022
Closest in time.