Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoEs) can scale up beyond traditional deep learning models by employing a routing strategy in which each input is processed by a single "expert" deep learning model.
Ueber die kleinste kugel, die eine räumliche figur einschliesst
Jung, H · 1901
Earlier work this paper cites.
Nouvelles applications des paramètres continus à la théorie des formes quadratiques. deuxième mémoire. recherches sur les parallélloèdres primitifs
Voronoi, G · 1908
Earlier work this paper cites.
On the mixture of distributions
Teicher, H · 1960
Earlier work this paper cites.
Identifiability of finite mixtures
Teicher, H · 1963
Earlier work this paper cites.
Systems of extremal control
Rastrigin, L. A · 1974
Earlier work this paper cites.
Colloquium lectures on geometric measure theory
Federer, H · 1978
Earlier work this paper cites.
A Connectionist Machine for Genetic Hillclimbing
Ackley, D · 1987
Earlier work this paper cites.
Learnability and the Vapnik-Chervonenkis dimension
Blumer, A., Ehrenfeucht, A., Haussler, D., and Warmuth, M. K · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Empirical processes: theory and applications , volume 2 of NSF-CBMS Regional Conference Series in Probability and Statistics
Pollard, D · 1990
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R · 1993
Earlier work this paper cites.
Optimal rate of convergence for finite mixture models
Chen, J · 1995
Earlier work this paper cites.
Convergence results for the em approach to mixtures of experts architectures
Jordan, M. I. and Xu, L · 1995
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
Mhaskar, H. N · 1996
Earlier work this paper cites.
Mixed poisson regression models with covariate dependent rates
Wang, P., Puterman, M. L., Cockburn, I., and Le, N · 1996
Earlier work this paper cites.
Geometric nonlinear functional analysis. Vol. 1 , volume 48 of American Mathematical Society Colloquium Publications
Benyamini, Y. and Lindenstrauss, J · 2000
Earlier work this paper cites.
Learning with mixtures of trees
Meila, M. and Jordan, M. I · 2000
Earlier work this paper cites.
Lectures on analysis on metric spaces
Heinonen, J · 2001
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Measured descent: a new embedding method for finite metrics
Krauthgamer, R., Lee, J. R., Mendel, M., and Naor, A · 2005
Earlier work this paper cites.
Exploring strategies for training deep neural networks
Larochelle, H., Bengio, Y., Louradour, J., and Lamblin, P · 2009
Earlier work this paper cites.
Convolutional deep belief networks on cifar-10
Krizhevsky, A. and Hinton, G · 2010
Earlier work this paper cites.
Dimensions, embeddings, and attractors , volume 186 of Cambridge Tracts in Mathematics
Robinson, J. C · 2011
Earlier work this paper cites.
Metric learning for large scale image classification: Generalizing to new classes at near-zero cost
Mensink, T., Verbeek, J., Perronnin, F., and Csurka, G · 2012
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Van Gool, L · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
The optimal sample complexity of PAC learning
Hanneke, S · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Bach, F · 2017
Earlier work this paper cites.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Harvey, N., Liaw, C., and Mehrabian, A · 2017
Earlier work this paper cites.
The expressive power of neural networks: A view from the width
Lu, Z., Pu, H., Wang, F., Hu, Z., and Wang, L · 2017
Earlier work this paper cites.
When and why are deep networks better than shallow ones?
Mhaskar, H., Liao, Q., and Poggio, T · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R · 2017
Cited alongside, same era.
SkipNet: Learning Dynamic Routing in Convolutional Networks
Wang, X., Yu, F., Dou, Z.-Y., Darrell, T., and Gonzalez, J. E · 2017
Cited alongside, same era.
Error bounds for approximations with deep relu networks
Yarotsky, D · 2017
Cited alongside, same era.
Universal discrete-time reservoir computers with stochastic inputs and linear readouts using non-homogeneous state-affine systems
Grigoryeva, L. and Ortega, J.-P · 2018
Cited alongside, same era.
Nonlinear approximation and (deep) relu networks
Daubechies, I., DeVore, R., Foucart, S., Hanin, B., and Petrova, G · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Later among the works it cites.
Galimberti, L., Kratsios, A., and Livieri, G · 2022
Later among the works it cites.
Convergence rates for gaussian mixtures of experts
Ho, N., Yang, C.-Y., and Jordan, M. I · 2022
Later among the works it cites.
Metric entropy limits on recurrent neural network learning of linear dynamical systems
Hutter, C., Gül, R., and Bölcskei, H · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Cited alongside, same era.
Optimal approximation of continuous functions by very deep relu networks
Yarotsky, D · 2018
Cited alongside, same era.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Bartlett, P. L., Harvey, N., Liaw, C., and Mehrabian, A · 2019
Cited alongside, same era.
Optimal approximation with sparsely connected deep neural networks
Bolcskei, H., Grohs, P., Kutyniok, G., and Petersen, P · 2019
Cited alongside, same era.
Dimensionality reduction for representing the knowledge of probabilistic models
Law, M. T., Snell, J., massoud Farahmand, A., Urtasun, R., and Zemel, R. S · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Adaptivity of deep reLU network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality
Suzuki, T · 2019
Cited alongside, same era.
Kratsios, A. and Papon, L · 2022
Later among the works it cites.
Do relu networks have an edge when approximating compactly-supported functions?
Kratsios, A. and Zamanlooy, B · 2022
Later among the works it cites.
Exponential relu dnn expression of holomorphic maps in high dimension
Opschoor, J. A., Schwab, C., and Zech, J · 2022
Later among the works it cites.
On the adversarial robustness of mixture of experts
Puigcerver, J., Jenatton, R., Riquelme, C., Awasthi, P., and Bhojanapalli, S · 2022
Later among the works it cites.
Expertnet: A symbiosis of classification and clustering, 2022
Srivastava, S., Kawaguchi, K., and Rajan, V · 2022
Later among the works it cites.
Hierarchical partition of unity networks: fast multilevel training
Trask, N., Henriksen, A., Martinez, C., and Cyr, E · 2022
Later among the works it cites.
Universal approximations of invariant maps by neural networks
Yarotsky, D · 2022
Later among the works it cites.
Learning sub-patterns in piecewise continuous functions
Zamanlooy, B. and Kratsios, A · 2022
Later among the works it cites.
Deep network approximation: Achieving arbitrary accuracy with fixed number of neurons
Zhang, S., Shen, Z., and Yang, H · 2022
Later among the works it cites.
Designing universal causal deep learning models: The geometric (hyper) transformer
Acciaio, B., Kratsios, A., and Pammer, G · 2023
Later among the works it cites.
Approximation theory of tree tensor networks: tensorized univariate functions
Ali, M. and Nouy, A · 2023
Later among the works it cites.
Neural networks in fréchet spaces
Benth, F. E., Detering, N., and Galimberti, L · 2023
Later among the works it cites.
Global universal approximation of functional input maps on weighted spaces
Cuchiero, C., Schmocker, P., and Teichmann, J · 2023
Later among the works it cites.
Attention enables zero approximation error, 2023
Fang, Z., Ouyang, Y., Zhou, D.-X., and Cheng, G · 2023
Later among the works it cites.
Approximation bounds for random neural networks and reservoir systems
Gonon, L., Grigoryeva, L., and Ortega, J.-P · 2023
Later among the works it cites.
Minimal width for universal property of deep rnn
hoon Song, C., Hwang, G., ho Lee, J., and Kang, M · 2023
Later among the works it cites.
Deep neural networks with ReLU-sine-exponential activations break curse of dimensionality in approximation on Hölder class
Jiao, Y., Lai, Y., Lu, X., Wang, F., Yang, J. Z., and Yang, Y · 2023
Later among the works it cites.
Strain-minimizing hyperbolic network embeddings with landmarks
Keller-Ressel, M. and Nargang, S · 2023
Later among the works it cites.
Rates of approximation by relu shallow neural networks
Mao, T. and Zhou, D.-X · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al · 2023
Later among the works it cites.
Lipschitz widths
Petrova, G. and Wojtaszczyk, P. a · 2023
Later among the works it cites.
The universal approximation theorem for complex-valued neural networks
Voigtlaender, F · 2023
Later among the works it cites.
Alternating gradient descent and mixture-of-experts for integrated multimodal perception
Akbari, H., Kondratyuk, D., Cui, Y., Hornung, R., Wang, H., and Adam, H · 2024
Closest in time.
Upper bound on vc-dimension of partitioned class
Conant, G · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Mixture of neural operators: Incorporating historical information for longer rollouts
Majid, H. A. and Tudisco, F · 2024
Closest in time.
Efficient learning using spiking neural networks equipped with affine encoders and decoders, 2024
Neuman, A. M. and Petersen, P. C · 2024
Closest in time.
Optimal rates of approximation by shallow reluk neural networks and applications to nonparametric regression
Yang, Y. and Zhou, D.-X · 2024
Closest in time.