Fetching the paper…
Reading the bibliography…
The sparsely gated mixture of experts (MoE) architecture sends different inputs to different subnetworks, i.e., experts, through trainable routers.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Earlier work this paper cites.
Thinet: A filter level pruning method for deep neural network compression
Luo, J.-H., Wu, J., and Lin, W · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S · 2017
Earlier work this paper cites.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Earlier work this paper cites.
Clip-q: Deep network compression learning by in-parallel pruning-quantization
Tung, F. and Mori, G · 2018
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained bert networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Earlier work this paper cites.
Learning parities with neural networks
Daniely, A. and Malach, E · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Guaranteed recovery of one-hidden-layer neural networks via cross entropy
Fu, H., Chi, Y., and Liang, Y · 2020
Cited alongside, same era.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Cited alongside, same era.
Efficient transformer-based large scale language representations using hardware-friendly block structured pruning
Li, B., Kong, Z., Zhang, T., Li, J., Li, Z., Liu, H., and Ding, C · 2020
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Later among the works it cites.
Generalization guarantee of training graph convolutional networks with graph topology sampling
Li, H., Wang, M., Liu, S., Chen, P.-Y., and Xiong, J · 2022
Later among the works it cites.
On the adversarial robustness of mixture of experts
Puigcerver, J., Jenatton, R., Riquelme, C., Awasthi, P., and Bhojanapalli, S · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Cited alongside, same era.
Computational separation between convolutional and fully-connected networks
Shalev-Shwartz, S. et al · 2020
Cited alongside, same era.
Structured pruning of large language models
Wang, Z., Wohlwend, J., and Lei, T · 2020
Cited alongside, same era.
An optimization and generalization analysis for max-pooling networks
Brutzkus, A. and Globerson, A · 2021
Cited alongside, same era.
Local signal adaptivity: Provable feature learning in neural networks beyond kernels
Karp, S., Winston, E., Li, Y., and Singh, A · 2021
Cited alongside, same era.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L · 2021
Cited alongside, same era.
Ebert: Efficient bert inference with dynamic structured pruning
Liu, Z., Li, F., Li, G., and Cheng, J · 2021
Cited alongside, same era.
Coca: Contrastive captioners are image-text foundation models
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y · 2022
Later among the works it cites.
Joint edge-model sparse learning is provably efficient for graph neural networks
Zhang, S., Wang, M., Chen, P.-Y., Liu, S., Lu, S., and Liu, M · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V. Y., Dai, A. M., Chen, Z., Le, Q. V., and Laudon, J · 2022
Later among the works it cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z. and Li, Y · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
Patch-level routing in mixture-of-experts is provably sample-efficient for convolutional neural networks
Chowdhury, M. N. R., Zhang, S., Wang, M., Liu, S., and Chen, P.-Y · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Later among the works it cites.
SparseGPT: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Later among the works it cites.
Instant soup: Cheap pruning ensembles in a single pass can draw lottery tickets from large models
Jaiswal, A. K., Liu, S., Chen, T., Ding, Y., and Wang, Z · 2023
Later among the works it cites.
Memory-efficient NLLB-200: Language-specific expert pruning of a massively multilingual machine translation model
Koishekenov, Y., Berard, A., and Nikoulina, V · 2023
Later among the works it cites.
A theoretical understanding of shallow vision transformers: Learning, generalization, and sample complexity
Li, H., Wang, M., Liu, S., and Chen, P.-Y · 2023
Later among the works it cites.
On the convergence and sample complexity analysis of deep q-networks with ϵ \epsilon -greedy exploration
Zhang, S., Li, H., Wang, M., Liu, M., Chen, P.-Y., Lu, S., Liu, S., Murugesan, K., and Chaudhury, S · 2023
Later among the works it cites.