Fetching the paper…
Reading the bibliography…
Human visual perception can easily generalize to out-of-distributed visual data, which is far beyond the capability of modern machine learning models.
Martín Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 1907
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vladimir Vapnik · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Michael I. Jordan and Robert A. Jacobs · 1994
Earlier work this paper cites.
A patient-adaptable ecg beat classifier using a mixture of experts approach
Yu Hen Hu, Surekha Palreddy, and Willis J Tompkins · 1997
Earlier work this paper cites.
Learning to perceive the world as articulated: an approach for hierarchical learning in sensory-motor systems
Jun Tani and Stefano Nolfi · 1999
Earlier work this paper cites.
Multi-head attention: Collaborate instead of concatenate
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2006
Earlier work this paper cites.
Learning visual attributes
Vittorio Ferrari and Andrew Zisserman · 2007
Earlier work this paper cites.
How neural networks extrapolate: From feedforward to graph neural networks
Keyulu Xu, Mozhi Zhang, Jingling Li, Simon S Du, Ken-ichi Kawarabayashi, and Stefanie Jegelka · 2009
Earlier work this paper cites.
Batch normalization embeddings for deep domain generalization
Mattia Segù, Alessio Tonioni, and Federico Tombari · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias
Chen Fang, Ye Xu, and Daniel N. Rockmore · 2013
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor S. Lempitsky · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep CORAL: correlation alignment for deep domain adaptation
Baochen Sun and Kate Saenko · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Deeper, broader and artier domain generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M. Hospedales · 2017
Earlier work this paper cites.
dsprites: Disentanglement testing sprites dataset
Loic Matthey, Irina Higgins, Demis Hassabis, and Alexander Lerchner · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep hashing network for unsupervised domain adaptation
Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, and Sethuraman Panchanathan · 2017
Earlier work this paper cites.
Random erasing data augmentation. arxiv
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang · 2017
Earlier work this paper cites.
Recognition in terra incognita
Sara Beery, Grant Van Horn, and Pietro Perona · 2018
Earlier work this paper cites.
Functional map of the world
Gordon A. Christie, Neil Fendley, James Wilson, and Ryan Mukherjee · 2018
Earlier work this paper cites.
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel · 2018
Earlier work this paper cites.
Multi-source domain adaptation with mixture of experts
Jiang Guo, Darsh J Shah, and Regina Barzilay · 2018
Cited alongside, same era.
Algorithms and theory for multiple-source adaptation
Judy Hoffman, Mehryar Mohri, and Ningshan Zhang · 2018
Cited alongside, same era.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Cited alongside, same era.
Domain generalization with adversarial feature learning
Haoliang Li, Sinno Jialin Pan, Shiqi Wang, and Alex C. Kot · 2018
Cited alongside, same era.
A unified feature disentangler for multi-domain image translation and manipulation
Alexander H. Liu, Yen-Cheng Liu, Yu-Ying Yeh, and Yu-Chiang Frank Wang · 2018
Cited alongside, same era.
Best sources forward: Domain generalization through source-specific nets
WILDS: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton Earnshaw, Imran Haque, Sara M. Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang · 2021
Later among the works it cites.
Out-of-distribution generalization via risk extrapolation (rex)
David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Rémi Le Priol, and Aaron C. Courville · 2021
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Later among the works it cites.
How does a neural network’s architecture impact its robustness to noisy labels?
Jingling Li, Mozhi Zhang, Keyulu Xu, John Dickerson, and Jimmy Ba · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Massimiliano Mancini, Samuel Rota Bulò, Barbara Caputo, and Elisa Ricci · 2018
Cited alongside, same era.
Domain generalization via model-agnostic learning of semantic features
Qi Dou, Daniel Coelho de Castro, Konstantinos Kamnitsas, and Ben Glocker · 2019
Cited alongside, same era.
Feature-critic networks for heterogeneous domain generalization
Yiying Li, Yongxin Yang, Wei Zhou, and Timothy Hospedales · 2019
Cited alongside, same era.
Moment matching for multi-source domain adaptation
Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Cited alongside, same era.
Learning to learn without forgetting by maximizing transfer and minimizing interference
Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, and Gerald Tesauro · 2019
Cited alongside, same era.
Fixing the train-test resolution discrepancy
Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Later among the works it cites.
Trivialaugment: Tuning-free yet state-of-the-art data augmentation
Samuel G Müller and Frank Hutter · 2021
Later among the works it cites.
Yao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan, Alex Beutel, and Xuezhi Wang · 2021
Later among the works it cites.
Dynamic inference with neural interpreters
Nasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter Gehler, Yoshua Bengio, Francesco Locatello, and Bernhard Schölkopf · 2021
Later among the works it cites.
Fishr: Invariant gradient variances for out-of-distribution generalization
Alexandre Rame, Corentin Dancette, and Matthieu Cord · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Later among the works it cites.
Gradient matching for domain generalization
Yuge Shi, Jeffrey Seely, Philip H. S. Torr, N. Siddharth, Awni Y. Hannun, Nicolas Usunier, and Gabriel Synnaeve · 2021
Later among the works it cites.
Reappraising domain generalization in neural networks
Sarath Sivaprasad, Akshay Goindani, Vaibhav Garg, and Vineet Gandhi · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Variational disentanglement for domain generalization
Yufei Wang, Haoliang Li, Lap-Pui Chau, and Alex C. Kot · 2021
Later among the works it cites.
A fine-grained analysis on distribution shift
Olivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre Alvise-Rebuffi, Ira Ktena, Taylan Cemgil, et al · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Later among the works it cites.
Towards a theoretical framework of out-of-distribution generalization
Haotian Ye, Chuanlong Xie, Tianle Cai, Ruichen Li, Zhenguo Li, and Liwei Wang · 2021
Later among the works it cites.
Domain generalization by mutual-information regularization with pre-trained models
Junbum Cha, Kyungjae Lee, Sungrae Park, and Sanghyuk Chun · 2022
Closest in time.
Towards understanding mixture of experts in deep learning
Zixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu, and Yuanzhi Li · 2022
Closest in time.
On the representation collapse of sparse mixture of experts
Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, and Furu Wei · 2022
Closest in time.
A review of sparse expert models in deep learning
William Fedus, Jeff Dean, and Barret Zoph · 2022
Closest in time.
Domain generalization using pretrained models without fine-tuning
Ziyue Li, Kan Ren, Xinyang Jiang, Bo Li, Haipeng Zhang, and Dongsheng Li · 2022
Closest in time.
On the principles of parsimony and self-consistency for the emergence of intelligence
Yi Ma, Doris Tsao, and Heung-Yeung Shum · 2022
Closest in time.
Multimodal contrastive learning with limoe: the language-image mixture of experts
Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, and Neil Houlsby · 2022
Closest in time.
How do vision transformers work?
Namuk Park and Songkuk Kim · 2022
Closest in time.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al · 2022
Closest in time.
Delving deep into the generalization of vision transformers under distribution shifts
Chongzhi Zhang, Mingyuan Zhang, Shanghang Zhang, Daisheng Jin, Qiang Zhou, Zhongang Cai, Haiyu Zhao, Xianglong Liu, and Ziwei Liu · 2022
Closest in time.
Sparse invariant risk minimization
Xiao Zhou, Yong Lin, Weizhong Zhang, and Tong Zhang · 2022
Closest in time.