Fetching the paper…
Reading the bibliography…
We provide a theoretical treatment of over-specified Gaussian mixtures of experts with covariate-free gating networks.
B. Yu · 1907
Earlier work this paper cites.
On the mixture of distributions
H. Teicher · 1960
Earlier work this paper cites.
Identifiability of mixtures
H. Teicher · 1961
Earlier work this paper cites.
Identifiability of finite mixtures
H. Teicher · 1963
Earlier work this paper cites.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
Optimal rate of convergence for finite mixture models
J. Chen · 1995
Earlier work this paper cites.
Convergence results for the EM approach to mixtures of experts architectures
M. I. Jordan and L. Xu · 1995
Earlier work this paper cites.
Mixture models: Theory, geometry and applications
B. Lindsay · 1995
Earlier work this paper cites.
Bayesian inference in mixtures-of-experts and hierarchical mixtures-of-experts models with an application to speech recognition
F. Peng, R. A. Jacobs, and M. A. Tanner · 1996
Earlier work this paper cites.
Mixed Poisson regression models with covariate dependent rates
P. Wang, M. L. Puterman, I. Cockburn, and N. Le · 1996
Earlier work this paper cites.
Empirical Processes in M-estimation
S. van de Geer · 2000
Cited alongside, same era.
Entropies and rates of convergence for maximum likelihood and bayes estimation for mixtures of normal densities
S. Ghosal and A. van der Vaart · 2001
Cited alongside, same era.
Infinite mixtures of Gaussian process experts
C. E. Rasmussen and Z. Ghahramani · 2002
Cited alongside, same era.
Topics in Optimal Transportation
Cédric Villani · 2003
Cited alongside, same era.
Variable selection in finite mixture of regression models
A. Khalili and J. Chen · 2007
Cited alongside, same era.
Asymptotic behaviour of the posterior distribution in overfitted mixture models
J. Rousseau and K. Mengersen · 2011
Cited alongside, same era.
Tensor decompositions for learning latent variable models
A. Anandkumar, R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky · 2015
Later among the works it cites.
Using mixtures in econometric models: a brief review and some new results
G. Compiani and Y. Kitamura · 2016
Later among the works it cites.
Convergence rates of parameter estimation for some weakly identifiable finite mixtures
N. Ho and X. Nguyen · 2016
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Later among the works it cites.
Strong identifiability and optimal minimax rates for finite mixture estimation
P. Heinrich and J. Kahn · 2018
Later among the works it cites.
Singularity structures and impacts on parameter estimation in finite mixtures of distributions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A method of moments for mixture models and hidden Markov models
A. Anandkumar, D. Hsu, and S. M. Kakade · 2012
Cited alongside, same era.
Mixture of regression models with varying mixing proportions: A semiparametric approach
M. Huang and W. Yao · 2012
Cited alongside, same era.
Nonparametric mixture of regression models
M. Huang, R. Li, and S. Wang · 2013
Cited alongside, same era.
Convergence of latent mixing measures in finite and infinite mixture models
X. Nguyen · 2013
Cited alongside, same era.
Learning factored representations in a deep mixture of experts
D. Eigen, M. Ranzato, and I. Sutskever · 2014
Cited alongside, same era.
Sharp analysis of expectation-maximization for weakly identifiable models
R. Dwivedi, N. Ho, K. Khamaru, M. J. Wainwright, M. I. Jordan, and B. Yu
Cited in the paper.
N. Ho and X. Nguyen · 2019
Closest in time.
Breaking the gridlock in mixture-of-experts: Consistent and efficient algorithms
A. Makkuva, P. Viswanath, S. Kannan, and S. Oh · 2019
Closest in time.
Learning in gated neural networks
A. Makkuva, P. Viswanath, S. Kannan, and S. Oh · 2020
Closest in time.
On the minimax optimality of the EM algorithm for learning two-component mixed linear regression
J. Y. Kwon, N. Ho, and C. Caramanis · 2021
Closest in time.
Towards statistical and computational complexities of Polyak step size gradient descent
T. Ren, F. Cui, A. Atsidakou, S. Sanghavi, and N. Ho · 2022
Closest in time.