Fetching the paper…
Reading the bibliography…
Mixture of experts (MoE), introduced over 20 years ago, is the simplest gated modular neural network architecture.
Task decomposition through competition in a modular connectionist architecture: The what and where vision tasks
Robert A. Jacobs, Michael I. Jordan, and Andrew G. Barto · 1991
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Lukasz Kaiser, Aidan N. Gomez, Noam Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit · 2017
Earlier work this paper cites.
Combining a Mixture of Experts with Transfer Learning in Complex Games
Dobre Mihai and Alex Lascarides · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Reinforcement Learning with Multiple Experts: A Bayesian Model Combination Approach
Michael Gimelfarb, Scott Sanner, and Chi-Guhn Lee · 2018
Cited alongside, same era.
Modular networks: Learning to decompose neural computation
Louis Kirsch, Julius Kunze, and David Barber · 2018
Cited alongside, same era.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi · 2018
Cited alongside, same era.
Att-MoE: Attention-based Mixture of Experts for nuclear and cytoplasmic segmentation
Jinhua Liu, Christian Desrosiers, and Yuanfeng Zhou · 2020
Cited alongside, same era.
Beyond distillation: Task-level mixture-of-experts for efficient inference
Sneha Kudugunta, Yanping Huang, Ankur Bapna, Maxim Krikun, Dmitry Lepikhin, Minh-Thang Luong, and Orhan Firat · 2021
Cited alongside, same era.
Efficient continual learning with modular networks and task-driven priors
Tom Veniat, Ludovic Denoyer, and Marc’Aurelio Ranzato · 2021
Later among the works it cites.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Later among the works it cites.
Mixture-of-variational-experts for continual learning
Heinke Hihn and Daniel Alexander Braun · 2022
Later among the works it cites.
Is a Modular Architecture Enough?
Sarthak Mittal, Yoshua Bengio, and Guillaume Lajoie · 2022
Later among the works it cites.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Cited alongside, same era.
Base layers: Simplifying training of large, sparse models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Cited alongside, same era.
Scaling Vision with Sparse Mixture of Experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Cited alongside, same era.
Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Minjia Zhang, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Jeff Rasley, Shaden Smith, Olatunji Ruwase, and Yuxiong He · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing, 2022
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew Dai, Zhifeng Chen, Quoc Le, and James Laudon · 2022
Later among the works it cites.