Fetching the paper…
Reading the bibliography…
Can models with particular structure avoid being biased towards spurious correlation in out-of-distribution (OOD) generalization? Peters et al.
Modular learning in neural networks
Ballard, D. H · 1987
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Fodor, J. A., Pylyshyn, Z. W., et al · 1988
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., Solla, S. A., Howard, R. E., and Jackel, L. D · 1989
Earlier work this paper cites.
Skeletonization: A Technique for Trimming the Fat from a Network via Relevance Assessment
Mozer, M. C. and Smolensky, P · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D. G · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Rethinking eliminative connectionism
Marcus, G. F · 1998
Earlier work this paper cites.
An overview of statistical learning theory
Vapnik, V. N · 1999
Earlier work this paper cites.
Modularity and community structure in networks
Newman, M. E · 2006
Earlier work this paper cites.
Risk variance penalization: From distributional robustness to causality
Xie, C., Chen, F., Liu, Y., and Li, Z · 2006
Earlier work this paper cites.
Learning from multiple sources
Crammer, K., Kearns, M., and Wortman, J · 2008
Earlier work this paper cites.
A theory of learning from different domains
Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W · 2010
Earlier work this paper cites.
On causal and anticausal learning
Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J · 2012
Earlier work this paper cites.
Xie, S. M., Kumar, A., Jones, R., Khani, F., Ma, T., and Liang, P · 2012
Earlier work this paper cites.
The evolutionary origins of modularity
Clune, J., Mouret, J.-B., and Lipson, H · 2013
Earlier work this paper cites.
Domain generalization via invariant feature representation
Muandet, K., Balduzzi, D., and Schölkopf, B · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T. Y., Maire, M., Belongie, S., Hays, J., and Zitnick, C. L · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Causal inference by using invariant prediction: identification and confidence intervals
Peters, J., Bühlmann, P., and Meinshausen, N · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., Chen, S., and Pan, S. J · 2017
Earlier work this paper cites.
Unified deep supervised domain adaptation and generalization
Motiian, S., Piccirilli, M., Adjeroh, D. A., and Doretto, G · 2017
Earlier work this paper cites.
Cognitive psychology for deep neural networks: A shape bias case study
Ritter, S., Barrett, D. G., Santoro, A., and Botvinick, M. M · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
Recognition in terra incognita
Beery, S., Van Horn, G., and Perona, P · 2018
Cited alongside, same era.
Automatically composing representation transformations as a means for generalization
Chang, M. B., Gupta, A., Levine, S., and Griffiths, T. L · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H · 2018
Neural networks are surprisingly modular
Filan, D., Hod, S., Wild, C., Critch, A., and Russell, S · 2020
Later among the works it cites.
Pruning neural networks at initialization: Why are we missing the mark?
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Later among the works it cites.
In search of lost domain generalization
Gulrajani, I. and Lopez-Paz, D · 2020
Later among the works it cites.
Characterising bias in compressed models
Hooker, S., Moorosi, N., Clark, G., Bengio, S., and Denton, E · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rethinking the value of network pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Cited alongside, same era.
Learning sparse neural networks through l0 regularization
Louizos, C., Welling, M., and Kingma, D. P · 2018
Cited alongside, same era.
Packnet: Adding multiple tasks to a single network by iterative pruning
Mallya, A. and Lazebnik, S · 2018
Cited alongside, same era.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Cited alongside, same era.
Invariant models for causal transfer learning
Rojas-Carulla, M., Schölkopf, B., Turner, R., and Peters, J · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
Mlprune: Multi-layer pruning for automated neural network compression
Zeng, W. and Urtasun, R · 2018
Cited alongside, same era.
Later among the works it cites.
Domain extrapolation via regret minimization
Jin, W., Barzilay, R., and Jaakkola, T · 2020
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Beery, S., et al · 2020
Later among the works it cites.
Out-of-distribution generalization with maximal invariant predictor
Koyama, M. and Yamaguchi, S · 2020
Later among the works it cites.
Out-of-distribution generalization via risk extrapolation (rex)
Krueger, D., Caballero, E., Jacobsen, J., Zhang, A., Binas, J., Priol, R. L., and Courville, A. C · 2020
Later among the works it cites.
Learning robust models using the principle of independent causal mechanisms
Müller, J., Schmier, R., Ardizzone, L., Rother, C., and Köthe, U · 2020
Later among the works it cites.
Understanding the failure modes of out-of-distribution generalization
Nagarajan, V., Andreassen, A., and Neyshabur, B · 2020
Later among the works it cites.
Learning from failure: Training debiased classifier from biased classifier
Nam, J., Cha, H., Ahn, S., Lee, J., and Shin, J · 2020
Later among the works it cites.
Learning explanations that are hard to vary
Parascandolo, G., Neitz, A., Orvieto, A., Gresele, L., and Schölkopf, B · 2020
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, S.-O., Bengio, Y., Courville, A., Precup, D., and Lajoie, G · 2020
Later among the works it cites.
An analysis of the adaptation speed of causal models
Priol, R. L., Harikandeh, R. B., Bengio, Y., and Lacoste-Julien, S · 2020
Later among the works it cites.
The risks of invariant risk minimization
Rosenfeld, E., Ravikumar, P., and Risteski, A · 2020
Later among the works it cites.
An investigation of why overparameterization exacerbates spurious correlations
Sagawa, S., Raghunathan, A., Koh, P. W., and Liang, P · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P · 2020
Later among the works it cites.
Informative dropout for robust representation learning: A shape-bias perspective
Shi, B., Zhang, D., Dai, Q., Zhu, Z., Mu, Y., and Wang, J · 2020
Later among the works it cites.
Sanity-checking pruning methods: Random tickets can win the jackpot
Su, J., Chen, Y., Cai, T., Wu, T., Gao, R., Wang, L., and Lee, J. D · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Wang, C., Zhang, G., and Grosse, R · 2020
Later among the works it cites.
Systematic generalisation with group invariant predictions
Ahmed, F., Bengio, Y., van Seijen, H., and Courville, A · 2021
Closest in time.
Empirical or invariant risk minimization? a sample complexity perspective
Ahuja, K., Wang, J., Dhurandhar, A., Shanmugam, K., and Varshney, K. R · 2021
Closest in time.
Recurrent independent mechanisms
Goyal, A., Lamb, A., Hoffmann, J., Sodhani, S., Levine, S., Bengio, Y., and Schölkopf, B · 2021
Closest in time.
Does invariant risk minimization capture invariance?
Kamath, P., Tangella, A., Sutherland, D. J., and Srebro, N · 2021
Closest in time.
Removing spurious features can hurt accuracy and affect groups disproportionately
Khani, F. and Liang, P · 2021
Closest in time.
Shape-texture debiased neural network training
Li, Y., Yu, Q., Tan, M., Mei, J., Tang, P., Shen, W., Yuille, A., and cihang xie · 2021
Closest in time.
Counterfactual generative networks
Sauer, A. and Geiger, A · 2021
Closest in time.