Fetching the paper…
Reading the bibliography…
Neural networks trained with SGD were recently shown to rely preferentially on linearly-predictive features and can ignore complex, equally-predictive ones.
The need for biases in learning generalizations
Tom M Mitchell · 1980
Earlier work this paper cites.
Improving generalization performance using double backpropagation
Harris Drucker and Yann LeCun · 1992
Earlier work this paper cites.
The supervised learning no-free-lunch theorems
David H Wolpert · 2002
Earlier work this paper cites.
An introduction to variable and feature selection
Isabelle Guyon and André Elisseeff · 2003
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba, Alexei A Efros, et al · 2011
Earlier work this paper cites.
Diversity regularized machine
Yang Yu, Yu-Feng Li, and Zhi-Hua Zhou · 2011
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Domain generalization for object recognition with multi-task autoencoders
Muhammad Ghifary, W Bastiaan Kleijn, Mengjie Zhang, and David Balduzzi · 2015
Earlier work this paper cites.
Diversity networks: Neural network compression using determinantal point processes
Zelda Mariet and Suvrit Sra · 2015
Earlier work this paper cites.
Pengtao Xie, Yuntian Deng, and Eric Xing · 2015
Earlier work this paper cites.
The multiverse loss for robust transfer learning
Etai Littwin and Lior Wolf · 2016
Earlier work this paper cites.
Causal inference by using invariant prediction: identification and confidence intervals
Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen · 2016
Earlier work this paper cites.
Situation recognition: Visual semantic role labeling for image understanding
Mark Yatskar, Luke Zettlemoyer, and Ali Farhadi · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Training ensembles to detect adversarial examples
Alexander Bagnall, Razvan Bunescu, and Gordon Stewart · 2017
Earlier work this paper cites.
Conditional variance penalties and domain shift robustness
Christina Heinze-Deml and Nicolai Meinshausen · 2017
Earlier work this paper cites.
Nonlinear ica of temporally dependent stationary sources
Aapo Hyvarinen and Hiroshi Morioka · 2017
Earlier work this paper cites.
Deeper, broader and artier domain generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales · 2017
Earlier work this paper cites.
Discovering causal signals in images
David Lopez-Paz, Robert Nishihara, Soumith Chintala, Bernhard Scholkopf, and Léon Bottou · 2017
Earlier work this paper cites.
Right for the right reasons: Training differentiable models by constraining their explanations
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Visual question answering: A tutorial
Damien Teney, Qi Wu, and Anton van den Hengel · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Deep convolutional networks do not classify based on global object shape
Nicholas Baker, Hongjing Lu, Gennady Erlikhman, and Philip J Kellman · 2018
Earlier work this paper cites.
Adapting auxiliary losses using gradient similarity
Yunshu Du, Wojciech M Czarnecki, Siddhant M Jayakumar, Mehrdad Farajtabar, Razvan Pascanu, and Balaji Lakshminarayanan · 2018
Earlier work this paper cites.
Domain generalization with domain-specific aggregation modules
Antonio D’Innocente and Barbara Caputo · 2018
Earlier work this paper cites.
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel · 2018
Earlier work this paper cites.
Learning qualitatively diverse and interpretable rules for classification
Andrew Slavin Ross, Weiwei Pan, and Finale Doshi-Velez · 2018
Cited alongside, same era.
Generalizing to unseen domains via adversarial data augmentation
Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John Duchi, Vittorio Murino, and Silvio Savarese · 2018
Cited alongside, same era.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Cited alongside, same era.
Rubi: Reducing unimodal biases in visual question answering
Remi Cadene, Corentin Dancette, Hedi Ben-younes, Matthieu Cord, and Devi Parikh · 2019
Cited alongside, same era.
Domain generalization by solving jigsaw puzzles
Fabio M Carlucci, Antonio D’Innocente, Silvia Bucci, Barbara Caputo, and Tatiana Tommasi · 2019
What shapes feature representations? exploring datasets, architectures, and training
Katherine L Hermann and Andrew K Lampinen · 2020
Later among the works it cites.
Overcoming language priors in VQA via decomposed linguistic representations
Chenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia, and Qi Wu · 2020
Later among the works it cites.
Towards nonlinear disentanglement in natural data with temporal sparse coding
David Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov, Wieland Brendel, Matthias Bethge, and Dylan Paiton · 2020
Later among the works it cites.
Out-of-distribution generalization via risk extrapolation (rex)
David Krueger, Ethan Caballero, Joern-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Remi Le Priol, and Aaron Courville · 2020
Later among the works it cites.
Domain generalization using a mixture of multiple latent domains
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer · 2019
Cited alongside, same era.
We need to talk about standard splits
Kyle Gorman and Steven Bedrick · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Cited alongside, same era.
Robust learning with jacobian regularization
Judy Hoffman, Daniel A Roberts, and Sho Yaida · 2019
Cited alongside, same era.
Scops: Self-supervised co-part segmentation
Wei-Chih Hung, Varun Jampani, Sifei Liu, Pavlo Molchanov, Ming-Hsuan Yang, and Jan Kautz · 2019
Cited alongside, same era.
Improving adversarial robustness of ensembles with diversity training
Sanjay Kariyappa and Moinuddin K Qureshi · 2019
Cited alongside, same era.
Unmasking clever hans predictors and assessing what machines really learn
Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder, Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller · 2019
Cited alongside, same era.
Toshihiko Matsuura and Tatsuya Harada · 2020
Later among the works it cites.
Dataless model selection with the deep frame potential
Calvin Murdock and Simon Lucey · 2020
Later among the works it cites.
Learning from failure: Training debiased classifier from biased classifier
Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin · 2020
Later among the works it cites.
Fairness through robustness: Investigating robustness disparity in deep learning
Vedant Nanda, Samuel Dooley, Sahil Singla, Soheil Feizi, and John P Dickerson · 2020
Later among the works it cites.
Learning diverse representations for fast adaptation to distribution shift
Daniel Pace, Alessandra Russo, and Murray Shanahan · 2020
Later among the works it cites.
Learning explanations that are hard to vary
Giambattista Parascandolo, Alexander Neitz, Antonio Orvieto, Luigi Gresele, and Bernhard Schölkopf · 2020
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Mohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron Courville, Doina Precup, and Guillaume Lajoie · 2020
Later among the works it cites.
Learning to learn single domain generalization
Fengchun Qiao, Long Zhao, and Xi Peng · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with checklist
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Later among the works it cites.
Ensembles of locally independent prediction models
Andrew Ross, Weiwei Pan, Leo Celi, and Finale Doshi-Velez · 2020
Later among the works it cites.
An investigation of why overparameterization exacerbates spurious correlations
Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh, and Percy Liang · 2020
Later among the works it cites.
Learning from others’ mistakes: Avoiding dataset biases without modeling them
Victor Sanh, Thomas Wolf, Yonatan Belinkov, and Alexander M Rush · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli · 2020
Later among the works it cites.
Avoiding the hypothesis-only bias in natural language inference via ensemble adversarial training
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Sebastian Riedel, and Tim Rocktäschel · 2020
Later among the works it cites.
Learning what makes a difference from counterfactual examples and gradient supervision
Damien Teney, Ehsan Abbasnedjad, and Anton van den Hengel · 2020
Later among the works it cites.
On the value of out-of-distribution testing: An example of goodhart’s law
Damien Teney, Kushal Kafle, Robik Shrestha, Ehsan Abbasnejad, Christopher Kanan, and Anton van den Hengel · 2020
Later among the works it cites.
Mind the trade-off: Debiasing nlu models without degrading the in-distribution performance
Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych · 2020
Later among the works it cites.
Towards debiasing nlu models from unknown biases
Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych · 2020
Later among the works it cites.
Noise or signal: The role of image backgrounds in object recognition
Kai Xiao, Logan Engstrom, Andrew Ilyas, and Aleksander Madry · 2020
Later among the works it cites.
Maximal multiverse learning for promoting cross-task generalization of fine-tuned language models
Itzik Malkiel and Lior Wolf · 2021
Closest in time.
A neural anisotropic view of underspecification in deep learning
Guillermo Ortiz-Jimenez, Itamar Franco Salazar-Reque, Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
Toward causal representation learning
Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio · 2021
Closest in time.
On calibration and out-of-domain generalization
Yoav Wald, Amir Feder, Daniel Greenfeld, and Uri Shalit · 2021
Closest in time.
ID and OOD performance are sometimes inversely correlated on real-world datasets
Damien Teney, Seong Joon Oh, and Ehsan Abbasnejad · 2022
Closest in time.
Predicting is not understanding: Recognizing and addressing underspecification in machine learning
Damien Teney, Maxime Peyrard, and Ehsan Abbasnejad · 2022
Closest in time.