Fetching the paper…
Reading the bibliography…
The Mixture-of-Experts (MoE) layer, a sparsely-activated model controlled by a router, has achieved great success in deep learning.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Bai, Y · 1910
Earlier work this paper cites.
Mixtures of linear regressions
De Veaux, R. D · 1989
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Jordan, M. I · 1994
Earlier work this paper cites.
Hidden markov decision trees
Jordan, M. I · 1997
Earlier work this paper cites.
Learning and approximation capabilities of adaptive spline activation function neural networks
Vecci, L · 1998
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Allen-Zhu, Z · 2001
Earlier work this paper cites.
Mixtures of gaussian processes
Tresp, V · 2001
Earlier work this paper cites.
A parallel mixture of svms for very large scale problems
Collobert, R · 2002
Earlier work this paper cites.
Conditional random fields for object recognition
Quattoni, A · 2004
Earlier work this paper cites.
Feature purification: How adversarial training performs robust deep learning
Allen-Zhu, Z · 2005
Earlier work this paper cites.
An end-to-end discriminative approach to machine translation
Liang, P · 2006
Earlier work this paper cites.
Variable selection in finite mixture of regression models
Khalili, A · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L · 2008
Earlier work this paper cites.
Twitter sentiment classification using distant supervision
Go, A · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Max-margin hidden conditional random fields for human action recognition
Wang, Y · 2009
Earlier work this paper cites.
Fitting mixtures of linear regressions
Faria, S · 2010
Cited alongside, same era.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z · 2012
Cited alongside, same era.
A method of moments for mixture models and hidden markov models
Anandkumar, A · 2012
Cited alongside, same era.
Identifiability and unmixing of latent parse trees
Hsu, D. J · 2012
Cited alongside, same era.
Spectral experts for estimating mixtures of linear regressions
Chaganty, A. T · 2013
Cited alongside, same era.
Learning factored representations in a deep mixture of experts
Eigen, D · 2013
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D · 2018
Later among the works it cites.
What can ResNet learn efficiently, going beyond kernels?
Allen-Zhu, Z · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S · 2019
Later among the works it cites.
Sentiment analysis of product reviews in russian using convolutional neural networks
Smetanin, S · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tensor decompositions for learning latent variable models
Anandkumar, A · 2014
Cited alongside, same era.
Alternating minimization for mixed linear regression
Yi, X · 2014
Cited alongside, same era.
High dimensional em algorithm: Statistical optimization and asymptotic normality
Wang, Z · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K · 2016
Cited alongside, same era.
Statistical guarantees for the em algorithm: From population to sample-based analysis
Balakrishnan, S · 2017
Cited alongside, same era.
Language modeling with gated convolutional networks
Dauphin, Y. N · 2017
Cited alongside, same era.
Later among the works it cites.
French sentiment analysis with bert
Blard, T · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond ntk
Li, Y · 2020
Later among the works it cites.
Tricks for training sparse translation models
Dua, D · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W · 2021
Later among the works it cites.
Adam is no better than normalized sgd: Dissecting how adaptivity improves gan performance
Jelassi, S · 2021
Later among the works it cites.
Base layers: Simplifying training of large, sparse models
Lewis, M · 2021
Later among the works it cites.
Hash layers for large sparse models
Roller, S · 2021
Later among the works it cites.
Toward understanding the feature learning process of self-supervised contrastive learning
Wen, Z · 2021
Later among the works it cites.
Understanding the generalization of adam in learning neural networks with proper regularization
Zou, D · 2021
Later among the works it cites.
Benign overfitting in two-layer convolutional neural networks
Cao, Y · 2022
Closest in time.