Fetching the paper…
Reading the bibliography…
Due to domain shifts, machine learning systems typically struggle to generalize well to new domains that differ from those of training data, which is what domain generalization (DG) aims to address.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Michael I. Jordan and Robert A. Jacobs · 1994
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vladimir Vapnik · 1999
Earlier work this paper cites.
Vicinal risk minimization
Olivier Chapelle, Jason Weston, Léon Bottou, and Vladimir Vapnik · 2000
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Domain adaptation via transfer component analysis
Sinno Jialin Pan, Ivor W. Tsang, James T. Kwok, and Qiang Yang · 2010
Earlier work this paper cites.
Undoing the damage of dataset bias
Aditya Khosla, Tinghui Zhou, Tomasz Malisiewicz, Alexei A. Efros, and Antonio Torralba · 2012
Earlier work this paper cites.
Twenty years of mixture of experts
Seniha Esen Yuksel, Joseph N. Wilson, and Paul D. Gader · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias
Chen Fang, Ye Xu, and Daniel N. Rockmore · 2013
Earlier work this paper cites.
Domain generalization via invariant feature representation
Krikamol Muandet, David Balduzzi, and Bernhard Schölkopf · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
Eric Tzeng, Judy Hoffman, Ning Zhang, Kate Saenko, and Trevor Darrell · 2014
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Yaroslav Ganin and Victor Lempitsky · 2015
Earlier work this paper cites.
Domain generalization for object recognition with multi-task autoencoders
Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang, and David Balduzzi · 2015
Earlier work this paper cites.
Learning attributes equals multi-source domain generalization
Chuang Gan, Tianbao Yang, and Boqing Gong · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
Scatter component analysis: A unified framework for domain adaptation and domain generalization
Muhammad Ghifary, David Balduzzi, W. Bastiaan Kleijn, and Mengjie Zhang · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V. Le · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Deep coral: Correlation alignment for deep domain adaptation
Baochen Sun and Kate Saenko · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Earlier work this paper cites.
Deeper, broader and artier domain generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M. Hospedales · 2017
Earlier work this paper cites.
Unified deep supervised domain adaptation and generalization
Saeid Motiian, Marco Piccirilli, Donald A. Adjeroh, and Gianfranco Doretto · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Deep hashing network for unsupervised domain adaptation
Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, and Sethuraman Panchanathan · 2017
Earlier work this paper cites.
Mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Earlier work this paper cites.
Metareg: Towards domain generalization using meta-regularization
Yogesh Balaji, Swami Sankaranarayanan, and Rama Chellappa · 2018
Earlier work this paper cites.
Recognition in terra incognita
Sara Beery, Grant Van Horn, and Pietro Perona · 2018
Cited alongside, same era.
Domain generalization with domain-specific aggregation modules
Antonio D’Innocente and Barbara Caputo · 2018
Cited alongside, same era.
Multi-source domain adaptation with mixture of experts
Jiang Guo, Darsh J. Shah, and Regina Barzilay · 2018
Cited alongside, same era.
A unified feature disentangler for multi-domain image translation and manipulation
Alexander H. Liu, Yen-Cheng Liu, Yu-Ying Yeh, and Yu-Chiang Frank Wang · 2018
Cited alongside, same era.
Stochastic hyperparameter optimization through hypernetworks
Jonathan Lorraine and David Duvenaud · 2018
Cited alongside, same era.
Improve unsupervised domain adaptation with mixup training
Shen Yan, Huan Song, Nanxiang Li, Lincan Zou, and Liu Ren · 2020
Later among the works it cites.
Meta-learning via hypernetworks
Dominic Zhao, Johannes von Oswald, Seijin Kobayashi, João Sacramento, and Benjamin F. Grewe · 2020
Later among the works it cites.
Learning to generate novel domains for domain generalization
Kaiyang Zhou, Yongxin Yang, Timothy Hospedales, and Tao Xiang · 2020
Later among the works it cites.
Invariance principle meets information bottleneck for out-of-distribution generalization
Kartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet, Yoshua Bengio, Ioannis Mitliagkas, and Irina Rish · 2021
Later among the works it cites.
Domain generalization by marginal transfer learning
Gilles Blanchard, Aniket Anand Deshmukh, Ürun Dogan, Gyemin Lee, and Clayton Scott · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Massimiliano Mancini, Samuel Rota Bulo, Barbara Caputo, and Elisa Ricci · 2018
Cited alongside, same era.
Generalizing across domains via cross-gradient training
Shiv Shankar, Vihari Piratla, Soumen Chakrabarti, Siddhartha Chaudhuri, Preethi Jyothi, and Sunita Sarawagi · 2018
Cited alongside, same era.
Generalizing to unseen domains via adversarial data augmentation
Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C. Duchi, Vittorio Murino, and Silvio Savarese · 2018
Cited alongside, same era.
Visual domain adaptation with manifold embedded distribution alignment
Jindong Wang, Wenjie Feng, Yiqiang Chen, Han Yu, Meiyu Huang, and Philip S. Yu · 2018
Cited alongside, same era.
Deep visual domain adaptation: A survey
Mei Wang and Weihong Deng · 2018
Cited alongside, same era.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Cited alongside, same era.
Domain generalization by solving jigsaw puzzles
Fabio M. Carlucci, Antonio D’Innocente, Silvia Bucci, Barbara Caputo, and Tatiana Tommasi · 2019
Cited alongside, same era.
Dhanajit Brahma, Vinay Kumar Verma, and Piyush Rai · 2021
Later among the works it cites.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning
Hussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy, Yihua Chen, Rahul Mazumder, Lichan Hong, and Ed Chi · 2021
Later among the works it cites.
Selfreg: Self-supervised contrastive regularization for domain generalization
Daehee Kim, Youngjun Yoo, Seunghyun Park, Jinkyu Kim, and Jaekoo Lee · 2021
Later among the works it cites.
Out-of-distribution generalization via risk extrapolation (rex)
David Krueger, Ethan Caballero, Joern-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Remi Le Priol, and Aaron Courville · 2021
Later among the works it cites.
A simple feature augmentation for domain generalization
Pan Li, Da Li, Wei Li, Shaogang Gong, Yanwei Fu, and Timothy M. Hospedales · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson · 2021
Later among the works it cites.
Reducing domain gap by reducing style bias
Hyeonseob Nam, HyunJae Lee, Jongchan Park, Wonjun Yoon, and Donggeun Yoo · 2021
Later among the works it cites.
Personalized federated learning using hypernetworks
Aviv Shamsian, Aviv Navon, Ethan Fetaya, and Gal Chechik · 2021
Later among the works it cites.
Hypergrid transformers: Towards a single model for multiple tasks
Yi Tay, Zhe Zhao, Dara Bahri, Don Metzler, and Da-Cheng Juan · 2021
Later among the works it cites.
Adaptive risk minimization: Learning to adapt to domain shift
Marvin Zhang, Henrik Marklund, Nikita Dhawan, Abhishek Gupta, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
Compound Domain Generalization via Meta-Knowledge Encoding
Chaoqi Chen, Jiongcheng Li, Xiaoguang Han, Xiaoqing Liu, and Yizhou Yu · 2022
Closest in time.
StableMoE: Stable routing strategy for mixture of experts
Damai Dai, Li Dong, Shuming Ma, Bo Zheng, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Closest in time.
Glam: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, and Orhan Firat · 2022
Closest in time.
A review of sparse expert models in deep learning
William Fedus, Jeff Dean, and Barret Zoph · 2022
Closest in time.
Sparse Fusion Mixture-of-Experts are Domain Generalizable Learners
Bo Li, Jingkang Yang, Jiawei Ren, Yezhen Wang, and Ziwei Liu · 2022
Closest in time.
Balancing Expert Utilization in Mixture-of-Experts Layers Embedded in CNNs
Svetlana Pavlitskaya, Christian Hubschneider, Lukas Struppek, and J. Marius Zöllner · 2022
Closest in time.
Fishr: Invariant gradient variances for out-of-distribution generalization
Alexandre Rame, Corentin Dancette, and Matthieu Cord · 2022
Closest in time.
Hypershot: Few-shot learning by kernel hypernetworks
Marcin Sendera, Marcin Przewięźlikowski, Konrad Karanowski, Maciej Zięba, Jacek Tabor, and Przemysław Spurek · 2022
Closest in time.
Example-based hypernetworks for out-of-distribution generalization
Tomer Volk, Eyal Ben-David, Ohad Amosy, Gal Chechik, and Roi Reichart · 2022
Closest in time.
Generalizing to unseen domains: A survey on domain generalization
Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Tao Qin, Wang Lu, Yiqiang Chen, Wenjun Zeng, and Philip Yu · 2022
Closest in time.
Towards principled disentanglement for domain generalization
Hanlin Zhang, Yi-Fan Zhang, Weiyang Liu, Adrian Weller, Bernhard Schölkopf, and Eric P Xing · 2022
Closest in time.
Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-Experts
Tao Zhong, Zhixiang Chi, Li Gu, Yang Wang, Yuanhao Yu, and Jin Tang · 2022
Closest in time.
Domain generalization: A survey
Kaiyang Zhou, Ziwei Liu, Yu Qiao, Tao Xiang, and Chen Change Loy · 2022
Closest in time.
Designing effective sparse expert models
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus · 2022
Closest in time.