Fetching the paper…
Reading the bibliography…
Sparse mixture of experts (SMoE) offers an appealing solution to scale up the model complexity beyond the mean of increasing the network's depth or width.
Participation of inhibitory and excitatory interneurones in the control of hippocampal cortical output
Andersen, P., Gross, G. N., Lomo, T., and Sveen, O · 1969
Earlier work this paper cites.
Interneuronal mechanisms in the cortex
Stefanis, C · 1969
Earlier work this paper cites.
Self-organization of orientation sensitive cells in the striate cortex
Von der Malsburg, C · 1973
Earlier work this paper cites.
Connectionist models and their properties
Feldman, J. A. and Ballard, D. H · 1982
Earlier work this paper cites.
Contour enhancement, short term memory, and constancies in reverberating neural networks
Grossberg, S. and Grossberg, S · 1982
Earlier work this paper cites.
Self-organized formation of topologically correct feature maps
Kohonen, T · 1982
Earlier work this paper cites.
Feature discovery by competitive learning
Rumelhart, D. E. and Zipser, D · 1985
Earlier work this paper cites.
Parallel distributed processing, volume 2: Explorations in the microstructure of cognition: Psychological and biological models , volume 2
McClelland, J. L., Rumelhart, D. E., Group, P. R., et al · 1987
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Jordan, M. I. and Jacobs, R. A · 1994
Earlier work this paper cites.
Hierarchical models of object recognition in cortex
Riesenhuber, M. and Poggio, T · 1999
Earlier work this paper cites.
Empirical Processes in M-estimation
van de Geer, S · 2000
Earlier work this paper cites.
Spiking inputs to a winner-take-all network
Oster, M. and Liu, S.-C · 2005
Earlier work this paper cites.
An l 1 l_{1} -oracle inequality for the Lasso in mixture-of-experts regression models
Nguyen, T., Nguyen, H. D., Chamroukhi, F., and McLachlan, G. J · 2009
Earlier work this paper cites.
New estimation and feature selection methods in mixture-of-experts models
Khalili, A · 2010
Earlier work this paper cites.
Approximation of conditional densities by smooth mixtures of regressions
Norets, A · 2010
Earlier work this paper cites.
Learning Word Vectors for Sentiment Analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Large text compression benchmark, 2011
Mahoney, M · 2011
Earlier work this paper cites.
On convergence rates of mixtures of polynomial experts
Mendes, E. F. and Jiang, W · 2012
Earlier work this paper cites.
Twenty Years of Mixture of Experts
Yuksel, S. E., Wilson, J. N., and Gader, P. D · 2012
Earlier work this paper cites.
Deep Learning of Representations: Looking Forward
Bengio, Y · 2013
Earlier work this paper cites.
The cerebellum as a neuronal machine
Eccles, J. C · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
Eigen, D., Ranzato, M., and Sutskever, I · 2013
Earlier work this paper cites.
Maxout networks
Goodfellow, I., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y · 2013
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Cited alongside, same era.
Compete to compute
Srivastava, R. K., Masci, J., Kazerounian, S., Gomez, F., and Schmidhuber, J · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Mixture of experts: a literature survey
Masoudnia, S. and Ebrahimpour, R · 2014
Cited alongside, same era.
Mixture of Gaussian regressions model with logistic weights, a penalized maximum likelihood approach
Montuelle, L. and Le Pennec, E · 2014
Cited alongside, same era.
A universal approximation theorem for mixture-of-experts models
Nguyen, H. D., Lloyd-Jones, L. R., and McLachlan, G. J · 2016
Model Selection and Approximation in High-dimensional Mixtures of Experts Models: from Theory to Practice
Nguyen, T · 2021
Later among the works it cites.
Approximation of probability density functions via location-scale finite mixtures in Lebesgue spaces
Nguyen, T., Chamroukhi, F., Nguyen, H. D., and McLachlan, G. J · 2021
Later among the works it cites.
Adaptive Bayesian estimation of conditional discrete-continuous distributions with an application to stock market trading activity
Norets, A. and Pelenis, J · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., and Houlsby, N · 2021
Later among the works it cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., and Fergus, R · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
HyperNetworks
Ha, D., Dai, A. M., and Le, Q. V · 2017
Cited alongside, same era.
Pointer Sentinel Mixture Models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Cited alongside, same era.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Å., and Polosukhin, I · 2017
Cited alongside, same era.
Practical and theoretical aspects of mixture-of-experts modeling: An overview
Nguyen, H. D. and Chamroukhi, F · 2018
Cited alongside, same era.
Approximation results regarding the multiple-output Gaussian gated mixture of linear experts model
Nguyen, H. D., Chamroukhi, F., and Forbes, F · 2019
Cited alongside, same era.
Scaling Vision with Sparse Mixture of Experts
Ruiz, C. R., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Pinto, A. S., Keysers, D., and Houlsby, N · 2021
Later among the works it cites.
Wang, Y., Wang, W., Joty, S., and Hoi, S. C · 2021
Later among the works it cites.
On the Representation Collapse of Sparse Mixture of Experts
Chi, Z., Dong, L., Huang, S., Dai, D., Ma, S., Patra, B., Singhal, S., Bajaj, P., Song, X., Mao, X.-L., Huang, H., and Wei, F · 2022
Later among the works it cites.
Unified Scaling Laws for Routed Language Models
Clark, A., De Las Casas, D., Guy, A., Mensch, A., Paganini, M., Hoffmann, J., Damoc, B., Hechtman, B., Cai, T., Borgeaud, S., Van Den Driessche, G. B., Rutherford, E., Hennigan, T., Johnson, M. J., Cassirer, A., Jones, C., Buchatskaya, E., Budden, D., Sifre, L., Osindero, S., Vinyals, O., Ranzato, M., Rae, J., Elsen, E., Kavukcuoglu, K., and Simonyan, K · 2022
Later among the works it cites.
StableMoE: Stable Routing Strategy for Mixture of Experts
Dai, D., Dong, L., Ma, S., Zheng, B., Sui, Z., Chang, B., and Wei, F · 2022
Later among the works it cites.
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M. P., Zhou, Z., Wang, T., Wang, E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q., Wu, Y., Chen, Z., and Cui, C · 2022
Later among the works it cites.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Later among the works it cites.
Convergence Rates for Gaussian Mixtures of Experts
Ho, N., Yang, C.-Y., and Jordan, M. I · 2022
Later among the works it cites.
Sparse Mixers: Combining MoE and Mixing to build a more efficient BERT
Lee-Thorp, J. and Ainslie, J · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Later among the works it cites.
A non-asymptotic approach for model selection via penalization in high-dimensional mixture of experts models
Nguyen, T., Nguyen, H. D., Chamroukhi, F., and Forbes, F · 2022
Later among the works it cites.
Adaptive Bayesian Estimation of Discrete-Continuous Distributions Under Smoothness and Sparsity
Norets, A. and Pelenis, J · 2022
Later among the works it cites.
Mixture-of-Experts with Expert Choice Routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Chen, z., Le, Q. V., and Laudon, J · 2022
Later among the works it cites.
Taming Sparsely Activated Transformer with Stochastic Experts
Zuo, S., Liu, X., Jiao, J., Kim, Y. J., Hassan, H., Zhang, R., Gao, J., and Zhao, T · 2022
Later among the works it cites.
Sparse moe as the new dropout: Scaling dense and self-slimmable transformers
Chen, T., Zhang, Z., JAISWAL, A. K., Liu, S., and Wang, Z · 2023
Later among the works it cites.
A Mixture-of-Expert Approach to RL-based Dialogue Management
Chow, Y., Tulepbergenov, A., Nachum, O., Gupta, D., Ryu, M., Ghavamzadeh, M., and Boutilier, C · 2023
Later among the works it cites.
HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts
Do, T., Khiem, L., Pham, Q., Nguyen, T., Doan, T.-N., Nguyen, B., Liu, C., Ramasamy, S., Li, X., and Hoi, S · 2023
Later among the works it cites.
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Later among the works it cites.
Brainformers: Trading simplicity for efficiency
Zhou, Y., Du, N., Huang, Y., Peng, D., Lan, C., Huang, D., Shakeri, S., So, D., Dai, A. M., Lu, Y., et al · 2023
Later among the works it cites.