Fetching the paper…
Reading the bibliography…
Sparsely-gated Mixture of Experts networks (MoEs) have demonstrated excellent scalability in Natural Language Processing.
Classification and regression trees
L. Breiman, J. Friedman, C. J. Stone, and R. A. Olshen · 1984
Earlier work this paper cites.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
A patient-adaptable ECG beat classifier using a mixture of experts approach
Y. H. Hu, S. Palreddy, and W. J. Tompkins · 1997
Earlier work this paper cites.
Time series prediction using mixtures of experts
A. J. Zeevi, R. Meir, and R. J. Adler · 1997
Earlier work this paper cites.
Improved learning algorithms for mixture of experts in multiclass classification
K. Chen, L. Xu, and H. Chi · 1999
Earlier work this paper cites.
Combining predictors: comparison of five meta machine learning methods
J. V. Hansen · 1999
Earlier work this paper cites.
Learning to perceive the world as articulated: an approach for hierarchical learning in sensory-motor systems
J. Tani and S. Nolfi · 1999
Earlier work this paper cites.
Learning methods for generic object recognition with invariance to pose and lighting
Y. LeCun, F. J. Huang, and L. Bottou · 2004
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories
F.-F. Li, R. Fergus, and P. Perona · 2004
Earlier work this paper cites.
Learning to reconstruct 3D human motion from Bayesian mixtures of experts. A probabilistic discriminative approach
C. Sminchisescu, A. Kanaujia, Z. Li, and D. Metaxas · 2004
Earlier work this paper cites.
Multi-cue pedestrian detection and tracking from a moving vehicle
D. M. Gavrila and S. Munder · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, and A. Y. Ng · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? The KITTI vision benchmark suite
A. Geiger, P. Lenz, and R. Urtasun · 2012
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar · 2012
Earlier work this paper cites.
Twenty years of mixture of experts
S. E. Yuksel, J. N. Wilson, and P. D. Gader · 2012
Earlier work this paper cites.
Deep learning of representations: Looking forward
Y. Bengio · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Low-rank approximations for conditional feedforward computation in deep neural networks
A. Davis and I. Arel · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
D. Eigen, M. Ranzato, and I. Sutskever · 2013
Earlier work this paper cites.
K. Cho and Y. Bengio · 2014
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Cited alongside, same era.
Deep sequential neural network
L. Denoyer and P. Gallinari · 2014
Cited alongside, same era.
Conditional computation in neural networks for faster models
E. Bengio, P.-L. Bacon, J. Pineau, and D. Precup · 2015
Cited alongside, same era.
Kaggle diabetic retinopathy detection, 2015
Kaggle and EyePacs · 2015
Cited alongside, same era.
Network of experts for large-scale image categorization
K. Ahmed, M. H. Baig, and L. Torresani · 2016
Cited alongside, same era.
Condconv: Conditionally parameterized convolutions for efficient inference
B. Yang, G. Bender, Q. V. Le, and J. Ngiam · 2019
Later among the works it cites.
A large-scale study of representation learning with the visual task adaptation benchmark
X. Zhai, J. Puigcerver, A. Kolesnikov, P. Ruyssen, C. Riquelme, M. Lucic, J. Djolonga, A. S. Pinto, M. Neumann, A. Dosovitskiy, L. Beyer, O. Bachem, M. Tschannen, M. Michalski, O. Bousquet, S. Gelly, and N. Houlsby · 2019
Later among the works it cites.
A large-scale study of representation learning with the visual task adaptation benchmark
X. Zhai, J. Puigcerver, A. Kolesnikov, P. Ruyssen, C. Riquelme, M. Lucic, J. Djolonga, A. S. Pinto, M. Neumann, A. Dosovitskiy, et al · 2019
Later among the works it cites.
Biased mixtures of experts: Enabling computer vision inference under data transfer limitations
A. Abbas and Y. Andreopoulos · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Hendrycks and K. Gimpel · 2016
Cited alongside, same era.
Remote sensing image scene classification: Benchmark and state of the art
G. Cheng, J. Han, and X. Lu · 2017
Cited alongside, same era.
Hard mixtures of experts for large scale weakly supervised vision
S. Gross, M. Ranzato, and A. Szlam · 2017
Cited alongside, same era.
The elements of statistical learning: data mining, inference, and prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2017
Cited alongside, same era.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, F.-F. Li, C. Lawrence Zitnick, and R. Girshick · 2017
Cited alongside, same era.
dSprites: Disentanglement testing sprites dataset, 2017
L. Matthey, I. Higgins, D. Hassabis, and A. Lerchner · 2017
Cited alongside, same era.
Routing networks: Adaptive selection of non-linear functions for multi-task learning
C. Rosenbaum, T. Klinger, and M. Riemer · 2017
Cited alongside, same era.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
E. D. Cubuk, B. Zoph, J. Shlens, and Q. Le · 2020
Later among the works it cites.
A baseline for few-shot image classification
G. S. Dhillon, P. Chaudhari, A. Ravichandran, and S. Soatto · 2020
Later among the works it cites.
Big transfer (BiT): General visual representation learning
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby · 2020
Later among the works it cites.
Using mixture of expert models to gain insights into semantic segmentation
S. Pavlitskaya, C. Hubschneider, M. Weber, R. Moritz, F. Huger, P. Schlicht, and M. Zollner · 2020
Later among the works it cites.
H. Pham, Z. Dai, Q. Xie, M.-T. Luong, and Q. V. Le · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2020
Later among the works it cites.
Deep mixture of experts via shallow embedding
X. Wang, F. Yu, L. Dunlap, Y.-A. Ma, R. Wang, A. Mirhoseini, T. Darrell, and J. E. Gonzalez · 2020
Later among the works it cites.
ViViT: A video vision transformer
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid · 2021
Closest in time.
High-performance large-scale image recognition without normalization
A. Brock, S. De, S. L. Smith, and K. Simonyan · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2021
Closest in time.
Token labeling: Training a 85.5% top-1 accuracy vision transformer with 56m parameters on imagenet
Z. Jiang, Q. Hou, L. Yuan, D. Zhou, X. Jin, A. Wang, and J. Feng · 2021
Closest in time.
GShard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2021
Closest in time.
Base layers: Simplifying training of large, sparse models
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer · 2021
Closest in time.
Carbon emissions and large neural network training
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean · 2021
Closest in time.
Going deeper with image transformers
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, F. E. Tay, J. Feng, and S. Yan · 2021
Closest in time.
Scaling vision transformers, 2021
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2021
Closest in time.