Fetching the paper…
Reading the bibliography…
In deep learning, mixture-of-experts (MoE) activates one or few experts (sub-networks) on a per-sample or per-token basis, resulting in significant computation reduction.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Jordan, M. I. and Jacobs, R. A · 1994
Earlier work this paper cites.
Improved learning algorithms for mixture of experts in multiclass classification
Chen, K., Xu, L., and Chi, H · 1999
Earlier work this paper cites.
Mixtures of gaussian processes
Tresp, V · 2000
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Allen-Zhu, Z. and Li, Y · 2001
Earlier work this paper cites.
A parallel mixture of SVMs for very large scale problems
Collobert, R., Bengio, S., and Bengio, Y · 2001
Earlier work this paper cites.
Infinite mixtures of gaussian process experts
Rasmussen, C. and Ghahramani, Z · 2001
Earlier work this paper cites.
Scaling large learning problems with hard parallel mixtures
Collobert, R., Bengio, Y., and Bengio, S · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
MNIST handwritten digit database. AT&T labs [online]. available http
LeCun, Y., Cortes, C., and Burges, C · 2010
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z. and Li, Y · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
Eigen, D., Ranzato, M., and Sutskever, I · 2013
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
Network of experts for large-scale image categorization
Ahmed, K., Baig, M. H., and Torresani, L · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Hard mixtures of experts for large scale weakly supervised vision
Gross, S., Ranzato, M., and Szlam, A · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
Diversity and depth in per-example routing models
Ramachandran, P. and Le, Q. V · 2018
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Allen-Zhu, Z. and Li, Y · 2019
When do neural networks outperform kernel methods?
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2020
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond NTK
Li, Y., Ma, T., and Zhang, H. R · 2020
Later among the works it cites.
Computational separation between convolutional and fully-connected networks
Shalev-Shwartz, S. et al · 2020
Later among the works it cites.
An optimization and generalization analysis for max-pooling networks
Brutzkus, A. and Globerson, A · 2021
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Bai, Y. and Lee, J. D · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.
Limitations of lazy training of two-layers neural network
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2019
Cited alongside, same era.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ji, Z. and Telgarsky, M · 2019
Cited alongside, same era.
Later among the works it cites.
Local signal adaptivity: Provable feature learning in neural networks beyond kernels
Karp, S., Winston, E., Li, Y., and Singh, A · 2021
Later among the works it cites.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L · 2021
Later among the works it cites.
Quantifying the benefit of using differentiable learning over tangent kernels
Malach, E., Kamath, P., Abbe, E., and Srebro, N · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., and Houlsby, N · 2021
Later among the works it cites.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Shi, Z., Wei, J., and Liang, Y · 2021
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2021
Later among the works it cites.
Feature purification: How adversarial training performs robust deep learning
Allen-Zhu, Z. and Li, Y · 2022
Later among the works it cites.
Towards understanding mixture of experts in deep learning
Chen, Z., Deng, Y., Wu, Y., Gu, Q., and Li, Y · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Later among the works it cites.
Learning and generalization of one-hidden-layer neural networks, going beyond standard gaussian data
Li, H., Zhang, S., and Wang, M · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V. Y., Dai, A. M., Chen, Z., Le, Q. V., and Laudon, J · 2022
Later among the works it cites.
A theoretical understanding of shallow vision transformers: Learning, generalization, and sample complexity
Li, H., Wang, M., Liu, S., and Chen, P.-Y · 2023
Closest in time.