Fetching the paper…
Reading the bibliography…
Neural networks trained with (stochastic) gradient descent have an inductive bias towards learning simpler solutions.
An analysis of the greedy algorithm for the submodular set covering problem
L. A. Wolsey · 1982
Earlier work this paper cites.
Silhouettes: a graphical aid to the interpretation and validation of cluster analysis
P. J. Rousseeuw · 1987
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Mathematical methods of statistics , volume 43
H. Cramér · 1999
Earlier work this paper cites.
Technical Report CNS-TR-2011-001, California Institute of Technology, 2011
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
Distributed submodular maximization: Identifying representative elements in massive data
B. Mirzasoleiman, A. Karbasi, R. Sarkar, and A. Krause · 2013
Earlier work this paper cites.
Streaming submodular maximization: Massive data summarization on the fly
A. Badanidiyuru, B. Mirzasoleiman, A. Karbasi, and A. Krause · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2014
Earlier work this paper cites.
Variance reduction in sgd by distributed importance sampling
G. Alain, A. Lamb, C. Sankar, A. Courville, and Y. Bengio · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Z. Liu, P. Luo, X. Wang, and X. Tang · 2015
Earlier work this paper cites.
Lazier than lazy greedy
B. Mirzasoleiman, A. Badanidiyuru, A. Karbasi, J. Vondrák, and A. Krause · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Towards principled methods for training generative adversarial networks
M. Arjovsky and L. Bottou · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2017
Cited alongside, same era.
Places: A 10 million image database for scene recognition
B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba · 2017
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Cited alongside, same era.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Cited alongside, same era.
Learning from failure: De-biasing classifier from biased classifier
J. Nam, H. Cha, S. Ahn, J. Lee, and J. Shin · 2020
Later among the works it cites.
An investigation of why overparameterization exacerbates spurious correlations
S. Sagawa, A. Raghunathan, P. W. Koh, and P. Liang · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
H. Shah, K. Tamuly, A. Raghunathan, P. Jain, and P. Netrapalli · 2020
Later among the works it cites.
No subclass left behind: Fine-grained robustness in coarse-grained classification problems
N. Sohoni, J. Dunnmon, G. Angus, A. Gu, and C. Ré · 2020
Later among the works it cites.
Environment inference for invariant learning
E. Creager, J.-H. Jacobsen, and R. Zemel · 2021
Later among the works it cites.
Just train twice: Improving group robustness without training group information
E. Z. Liu, B. Haghgoo, A. S. Chen, A. Raghunathan, P. W. Koh, S. Sagawa, P. Liang, and C. Finn · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Nakkiran, G. Kaplun, D. Kalimeris, T. Yang, B. L. Edelman, F. Zhang, and B. Barak · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Cited alongside, same era.
Distributionally robust neural networks
S. Sagawa, P. W. Koh, T. B. Hashimoto, and P. Liang · 2019
Cited alongside, same era.
Robustness may be at odds with accuracy
D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry · 2019
Cited alongside, same era.
Systematic generalisation with group invariant predictions
F. Ahmed, Y. Bengio, H. van Seijen, and A. Courville · 2020
Cited alongside, same era.
What shapes feature representations? exploring datasets, architectures, and training
K. Hermann and A. Lampinen · 2020
Cited alongside, same era.
The surprising simplicity of the early-time learning dynamics of neural networks
W. Hu, L. Xiao, B. Adlam, and J. Pennington · 2020
Cited alongside, same era.
Later among the works it cites.
Spread spurious attribute: Improving worst-group accuracy with spurious attribute estimation
J. Nam, J. Kim, J. Lee, and J. Shin · 2021
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
M. Pezeshki, O. Kaba, Y. Bengio, A. C. Courville, D. Precup, and G. Lajoie · 2021
Later among the works it cites.
Robust representation learning via perceptual similarity metrics
S. A. Taghanaki, K. Choi, A. H. Khasahmadi, and A. Goyal · 2021
Later among the works it cites.
Evading the simplicity bias: Training a diverse set of models discovers solutions with superior ood generalization
D. Teney, E. Abbasnejad, S. Lucey, and A. Van den Hengel · 2022
Later among the works it cites.
Correct-n-contrast: A contrastive approach for improving robustness to spurious correlations
M. Zhang, N. S. Sohoni, H. R. Zhang, C. Finn, and C. Ré · 2022
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
P. Kirichenko, P. Izmailov, and A. G. Wilson · 2023
Closest in time.
A whac-a-mole dilemma: Shortcuts come in multiples where mitigating one amplifies others
Z. Li, I. Evtimov, A. Gordo, C. Hazirbas, T. Hassner, C. C. Ferrer, C. Xu, and M. Ibrahim · 2023
Closest in time.