Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) models are a promising way to scale up model capacity without significantly increasing computational cost.
A Relationship Between Arbitrary Positive Matrices and Doubly Stochastic Matrices
Richard Sinkhorn · 1964
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Twenty years of mixture of experts
Seniha Esen Yuksel, Joseph N Wilson, and Paul D Gader · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Earlier work this paper cites.
Network of experts for large-scale image categorization
Karim Ahmed, Mohammad Haris Baig, and Lorenzo Torresani · 2016
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2016
Earlier work this paper cites.
Hard mixtures of experts for large scale weakly supervised vision
Sam Gross, Marc’Aurelio Ranzato, and Arthur Szlam · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Smooth and sparse optimal transport
Mathieu Blondel, Vivien Seguy, and Antoine Rolet · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi · 2018
Earlier work this paper cites.
Computational Optimal Transport: With Applications to Data Science
Gabriel Peyré and Marco Cuturi · 2019
Cited alongside, same era.
CondConv: Conditionally parameterized convolutions for efficient inference
Brandon Yang, Gabriel Bender, Quoc V Le, and Jiquan Ngiam · 2019
Cited alongside, same era.
Biased mixtures of experts: Enabling computer vision inference under data transfer limitations
Alhabib Abbas and Yiannis Andreopoulos · 2020
Cited alongside, same era.
Big transfer (BiT): general visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2020
Cited alongside, same era.
Using mixture of expert models to gain insights into semantic segmentation
Svetlana Pavlitskaya, Christian Hubschneider, Michael Weber, Ruby Moritz, Fabian Huger, Peter Schlicht, and Marius Zollner · 2020
Cited alongside, same era.
Deep mixture of experts via shallow embedding
Towards understanding the mixture-of-experts layer in deep learning
Zixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu, and Yuanzhi Li · 2022
Later among the works it cites.
Unified scaling laws for routed language models
Aidan Clark, Diego de las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al · 2022
Later among the works it cites.
Sparse mixture-of-experts are domain generalizable learners
Bo Li, Yifei Shen, Jingkang Yang, Yezhen Wang, Jiawei Ren, Tong Che, Jun Zhang, and Ziwei Liu · 2022
Later among the works it cites.
Multimodal contrastive learning with LIMoE: the language-image mixture of experts
Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, and Neil Houlsby · 2022
Later among the works it cites.
On the adversarial robustness of mixture of experts
Joan Puigcerver, Rodolphe Jenatton, Carlos Riquelme, Pranjal Awasthi, and Srinadh Bhojanapalli · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xin Wang, Fisher Yu, Lisa Dunlap, Yi-An Ma, Ruth Wang, Azalia Mirhoseini, Trevor Darrell, and Joseph E Gonzalez · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Unbiased gradient estimation with balanced assignments for mixtures of experts
Wouter Kool, Chris J. Maddison, and Andriy Mnih · 2021
Cited alongside, same era.
GShard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Cited alongside, same era.
Base layers: Simplifying training of large, sparse models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Cited alongside, same era.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Cited alongside, same era.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, and Jason E Weston · 2021
Cited alongside, same era.
Lemeng Wu, Mengchen Liu, Yinpeng Chen, Dongdong Chen, Xiyang Dai, and Lu Yuan · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Y Zhao, Andrew M. Dai, Zhifeng Chen, Quoc V Le, and James Laudon · 2022
Later among the works it cites.
Mod-squad: Designing mixtures of experts as modular multi-task learners
Zitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen, Hengshuang Zhao, Erik Learned-Miller, and Chuang Gan · 2023
Later among the works it cites.
Tutel: Adaptive mixture-of-experts at scale
Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, et al · 2023
Later among the works it cites.
Sparse upcycling: Training mixture-of-experts from dense checkpoints
Aran Komatsuzaki, Joan Puigcerver, James Lee-Thorp, Carlos Riquelme Ruiz, Basil Mustafa, Joshua Ainslie, Yi Tay, Mostafa Dehghani, and Neil Houlsby · 2023
Later among the works it cites.
Sparsity-constrained optimal transport
Tianlin Liu, Joan Puigcerver, and Mathieu Blondel · 2023
Later among the works it cites.
From sparse to soft mixtures of experts
Joan Puigcerver, Carlos Riquelme, Basil Mustafa, and Neil Houlsby · 2023
Later among the works it cites.
Fast, differentiable and sparse top-k: a convex analysis perspective
Michael Eli Sander, Joan Puigcerver, Josip Djolonga, Gabriel Peyré, and Mathieu Blondel · 2023
Later among the works it cites.