Fetching the paper…
Reading the bibliography…
Standard neural networks struggle to generalize under distribution shifts in computer vision.
Neural network ensembles
Lars Kai Hansen and Peter Salamon · 1990
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Bradley Efron · 1992
Earlier work this paper cites.
Neural network ensembles, cross validation, and active learning
Anders Krogh and Jesper Vedelsby · 1995
Earlier work this paper cites.
Generalization error of ensemble estimators
Naonori Ueda and Ryohei Nakano · 1996
Earlier work this paper cites.
Bias plus variance decomposition for zero-one loss functions
Ron Kohavi, David H Wolpert, et al · 1996
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
On the effect of data set size on bias and variance in classification learning
Damien Brain and Geoffrey I Webb · 1999
Earlier work this paper cites.
A unified bias-variance decomposition
Pedro Domingos · 2000
Earlier work this paper cites.
Ensemble methods in machine learning
Thomas Dietterich · 2000
Earlier work this paper cites.
Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy
Ludmila I Kuncheva and Christopher J Whitaker · 2003
Earlier work this paper cites.
Comparison of classifier selection methods for improving committee performance
Matti Aksela · 2003
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2003
Earlier work this paper cites.
Between two extremes: Examining decompositions of the ensemble objective function
Gavin Brown, Jeremy Wyatt, and Ping Sun · 2005
Earlier work this paper cites.
How to normalize a kernel matrix
Jason Rennie · 2005
Earlier work this paper cites.
Normalized kernels as similarity indices
Julien Ah-Pine · 2010
Earlier work this paper cites.
Mnist handwritten digit database, 2010
Yann LeCun, Corinna Cortes, and Chris Burges · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola · 2012
Earlier work this paper cites.
Domain generalization via invariant feature representation
Krikamol Muandet, David Balduzzi, and Bernhard Schölkopf · 2013
Earlier work this paper cites.
Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias
Chen Fang, Ye Xu, and Daniel N Rockmore · 2013
Earlier work this paper cites.
Gaussian processes for nonlinear signal processing: An overview of recent advances
Fernando Pérez-Cruz, Steven Van Vaerenbergh, Juan José Murillo-Fuentes, Miguel Lázaro-Gredilla, and Ignacio Santamaria · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Causal inference by using invariant prediction: identification and confidence intervals
Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen · 2016
Earlier work this paper cites.
Return of frustratingly easy domain adaptation
Baochen Sun, Jiashi Feng, and Kate Saenko · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of african-american english
Su Lin Blodgett, Lisa Green, and Brendan O’Connor · 2016
Earlier work this paper cites.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst · 2016
Earlier work this paper cites.
Swapout: Learning an ensemble of deep architectures
Saurabh Singh, Derek Hoiem, and David Forsyth · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Deep hashing network for unsupervised domain adaptation
Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, and Sethuraman Panchanathan · 2017
Earlier work this paper cites.
Deeper, broader and artier domain generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study
John R. Zech, Marcus A. Badgeley, Manway Liu, Anthony B. Costa, Joseph J. Titano, and Eric Karl Oermann · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Deep learning using rectified linear units (relu)
Abien Fred Agarap · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Recognition in Terra Incognita
Sara Beery, Grant Van Horn, and Pietro Perona · 2018
SWAD: Domain generalization by seeking flat minima
Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park · 2021
Later among the works it cites.
Ensemble of averages: Improving model selection and boosting performance in domain generalization
Devansh Arpit, Huan Wang, Yingbo Zhou, and Caiming Xiong · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Reproducing kernel hilbert space, mercer’s theorem, eigenfunctions, nystrom method, and use of kernels in machine learning: Tutorial and survey
Benyamin Ghojogh, Ali Ghodsi, Fakhri Karray, and Mark Crowley · 2021
Later among the works it cites.
WILDS: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton Earnshaw, Imran Haque, Sara M Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Cited alongside, same era.
Invariant risk minimization
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Cited alongside, same era.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek · 2019
Cited alongside, same era.
Deep ensembles: A loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton · 2019
Cited alongside, same era.
Later among the works it cites.
Multi-domain ensembles for domain generalization
Kowshik Thopalli, Sameeksha Katoch, Jayaraman J. Thiagarajan, Pavan K. Turaga, and Andreas Spanias · 2021
Later among the works it cites.
Robustness via cross-domain ensembles
Teresa Yeo, Oguzhan Fatih Kar, and Amir Roshan Zamir · 2021
Later among the works it cites.
Combining ensembles and data augmentation can harm your calibration
Yeming Wen, Ghassen Jerfel, Rafael Muller, Michael W Dusenberry, Jasper Snoek, Balaji Lakshminarayanan, and Dustin Tran · 2021
Later among the works it cites.
MixMo: Mixing multiple inputs for multiple outputs via deep subnetworks
Alexandre Rame, Remy Sun, and Matthieu Cord · 2021
Later among the works it cites.
DICE: Diversity in deep ensembles via conditional redundancy adversarial estimation
Alexandre Ramé and Matthieu Cord · 2021
Later among the works it cites.
Evading the simplicity bias: Training a diverse set of models discovers solutions with superior ood generalization
Damien Teney, Ehsan Abbasnejad, Simon Lucey, and Anton van den Hengel · 2021
Later among the works it cites.
Loss surface simplexes for mode connecting volumes and fast ensembling
Gregory Benton, Wesley Maddox, Sanae Lotfi, and Andrew Gordon Gordon Wilson · 2021
Later among the works it cites.
Learning neural network subspaces
Mitchell Wortsman, Maxwell Horton, Carlos Guestrin, Ali Farhadi, and Mohammad Rastegari · 2021
Later among the works it cites.
Relative flatness and generalization
Henning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu, and Mario Boley · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Fishr: Invariant gradient variances for out-of-distribution generalization
Alexandre Rame, Corentin Dancette, and Matthieu Cord · 2022
Closest in time.
Ood-bench: Benchmarking and understanding out-of-distribution generalization datasets and algorithms
Nanyang Ye, Kaican Li, Lanqing Hong, Haoyue Bai, Yiting Chen, Fengwei Zhou, and Zhenguo Li · 2022
Closest in time.
No one representation to rule them all: Overlapping features of training methods
Raphael Gontijo-Lopes, Yann Dauphin, and Ekin Dogus Cubuk · 2022
Closest in time.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Hanna Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt · 2022
Closest in time.
Merging models with Fisher-weighted averaging
Michael Matena and Colin Raffel · 2022
Closest in time.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt · 2022
Closest in time.
When do flat minima optimizers work?
Jean Kaddour, Linqing Liu, Ricardo Silva, and Matt Kusner · 2022
Closest in time.
Optimal representations for covariate shift
Yangjun Ruan, Yann Dubois, and Chris J. Maddison · 2022
Closest in time.
Neural tangent kernel beyond the infinite-width limit: Effects of depth and initialization
Mariia Seleznova and Gitta Kutyniok · 2022
Closest in time.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang · 2022
Closest in time.
Domain generalization using ensemble learning
Yusuf Mesbah, Youssef Youssry Ibrahim, and Adil Mehood Khan · 2022
Closest in time.
Domain generalization using pretrained models without fine-tuning
Ziyue Li, Kan Ren, Xinyang Jiang, Bo Li, Haipeng Zhang, and Dongsheng Li · 2022
Closest in time.
Diversify and disambiguate: Learning from underspecified data
Yoonho Lee, Huaxiu Yao, and Chelsea Finn · 2022
Closest in time.
Agree to disagree: Diversity through disagreement for better transferability
Matteo Pagliardini, Martin Jaggi, François Fleuret, and Sai Praneeth Karimireddy · 2022
Closest in time.
Stochastic weight averaging revisited
Hao Guo, Jiyong Jin, and Bin Liu · 2022
Closest in time.
Fusing finetuned models for better pretraining
Leshem Choshen, Elad Venezian, Noam Slonim, and Yoav Katz · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
Domain-adjusted regression or: Erm may already learn features sufficient for out-of-distribution generalization
Elan Rosenfeld, Pradeep Ravikumar, and Andrej Risteski · 2022
Closest in time.
Last layer re-training is sufficient for robustness to spurious correlations
Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson · 2023
Closest in time.