Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) is a widely popular model for ensemble learning and is a basic building block of highly successful modern neural networks as well as a component in Gated Recurrent Units (GRU) and Attention networks.
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
Stein, C · 1972
Earlier work this paper cites.
Three-way arrays: rank and uniqueness of trilinear decompositions, with application to arithmetic complexity and statistics
Kruskal, J. B · 1977
Earlier work this paper cites.
Airfoil self-noise and prediction
Brooks, T., Pope, D., and Marcolini., A · 1989
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Jordan, M. I. and Jacobs, R. A · 1994
Earlier work this paper cites.
Convergence results for the EM approach to mixtures of experts architectures
Jordan, M. I. and Xu, L · 1995
Earlier work this paper cites.
Modeling of strength of high performance concrete using artificial neural networks
Yeh, I.-C · 1998
Earlier work this paper cites.
Mixtures of gaussian processes
Tresp, V · 2001
Earlier work this paper cites.
A parallel mixture of SVMs for very large scale problems
Collobert, R., Bengio, S., and Bengio, Y · 2002
Earlier work this paper cites.
The EM algorithm and extensions , volume 382
McLachlan, G. and Krishnan, T · 2007
Earlier work this paper cites.
Twenty years of mixture of experts
Yuksel, S. E., Wilson, J. N., and Gader, P. D · 2012
Earlier work this paper cites.
Tensor decompositions for learning latent variable models
Anandkumar, A., Ge, R., Hsu, D., Kakade, S. M., and Telgarsky, M · 2014
Earlier work this paper cites.
A convex formulation for mixed regression with two components: Minimax optimal rates
Chen, Y., Yi, X., and Caramanis, C · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gülçehre, Ç., Cho, K., and Bengio, Y · 2014
Cited alongside, same era.
Score function features for discriminative learning: Matrix and tensor framework
Janzamin, M., Sedghi, H., and Anandkumar, A · 2014
Cited alongside, same era.
Mixture of experts: a literature survey
Masoudnia, S. and Ebrahimpour, R · 2014
Cited alongside, same era.
Hierarchical mixture-of-experts model for large-scale gaussian process regression
Ng, J. W. and Deisenroth, M. P · 2014
Cited alongside, same era.
Provable tensor methods for learning mixtures of classifiers
Sedghi, H., Janzamin, M., and Anandkumar, A · 2014
Statistical guarantees for the EM algorithm: From population to sample-based analysis
Balakrishnan, S., Wainwright, M. J., and Yu, B · 2017
Later among the works it cites.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
Du, S. S., Lee, J. D., Tian, Y., Poczos, B., and Singh, A · 2017
Later among the works it cites.
Learning one-hidden-layer neural networks with landscape design
Ge, R., Lee, J. D., and Ma, T · 2017
Later among the works it cites.
Hard mixtures of experts for large scale weakly supervised vision
Gross, S., Szlam, A., et al · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Li, Y. and Yuan, Y · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning mixtures of linear classifiers
Sun, Y., Ioannidis, S., and Montanari, A · 2014
Cited alongside, same era.
Alternating minimization for mixed linear regression
Yi, X., Caramanis, C., and Sanghavi, S · 2014
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Janzamin, M., Sedghi, H., and Anandkumar, A · 2015
Cited alongside, same era.
Generative image modeling using spatial lstms
Theis, L. and Bethge, M · 2015
Cited alongside, same era.
Jin, C., Zhang, Y., Balakrishnan, S., Wainwright, M. J., and Jordan, M · 2016
Cited alongside, same era.
Lstm-based mixture-of-experts for knowledge-aware dialogues
Le, P., Dymetman, M., and Renders, J.-M · 2016
Cited alongside, same era.
Yi, X., Caramanis, C., and Sanghavi, S · 2016
Cited alongside, same era.
Safran, I. and Shamir, O · 2017
Later among the works it cites.
Orthogonalized ALS: A theoretically principled tensor decomposition algorithm for practical use
Sharan, V. and Valiant, G · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Later among the works it cites.
Human-machine conversation based on hybrid neural network
Sun, X., Peng, X., Ren, F., and Xue, Y · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S · 2017
Later among the works it cites.
Deep mixture of experts via shallow embedding
Wang, X., Yu, F., Wang, R., Ma, Y.-A., Mirhoseini, A., Darrell, T., and Gonzalez, J. E · 2018
Closest in time.
Using mixture design and neural networks to build stock selection decision support systems
Liu, Y.-C. and Yeh, I.-C · 2090
Closest in time.