Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) architectures offer a general solution to the high inference costs of large language models (LLMs) via sparse routing, bringing faster and more accurate models, at the cost of massive parameter counts.
A method for the construction of minimum-redundancy codes
Huffman, D. A · 1952
Earlier work this paper cites.
A technique for high-performance data compression
Welch, T. A · 1984
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J · 2003
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Fast sparse convnets
Elsen, E., Dukhan, M., Gale, T., and Simonyan, K · 2020
Earlier work this paper cites.
Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Gshard, Z · 2020
Earlier work this paper cites.
Up or down? Adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Earlier work this paper cites.
Towards accurate post-training network quantization via bit-split and stitching
Wang, P., Chen, Q., He, X., and Cheng, J · 2020
Earlier work this paper cites.
A survey of quantization methods for efficient neural network inference
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Earlier work this paper cites.
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Earlier work this paper cites.
Accurate post training quantization with small calibration sets
Hubara, I., Nahshan, Y., Hanani, Y., Banner, R., and Soudry, D · 2021
Earlier work this paper cites.
A white paper on neural network quantization
Nagel, M., Fournarakis, M., Amjad, R. A., Bondarenko, Y., van Baalen, M., and Blankevoort, T · 2021
Earlier work this paper cites.
Hash layers for large sparse models
Roller, S., Sukhbaatar, S., Weston, J., et al · 2021
Cited alongside, same era.
Efficient large scale language modeling with mixtures of experts
Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X. V., Du, J., Iyer, S., Pasunuru, R., et al · 2022
Cited alongside, same era.
Pathways: Asynchronous distributed dataflow for ml
Barham, P., Chowdhery, A., Dean, J., Ghemawat, S., Hand, S., Hurt, D., Isard, M., Lim, H., Pang, R., Roy, S., et al · 2022
Cited alongside, same era.
Task-specific expert pruning for sparse mixture-of-experts
Chen, T., Huang, S., Xie, Y., Jiao, B., Jiang, D., Zhou, H., Li, J., and Wei, F · 2022
Cited alongside, same era.
Unified scaling laws for routed language models
Clark, A., De Las Casas, D., Guy, A., Mensch, A., Paganini, M., Hoffmann, J., Damoc, B., Hechtman, B., Cai, T., Borgeaud, S., et al · 2022
Cited alongside, same era.
ST-MoE: Designing stable and transferable sparse expert models
Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W · 2022
Later among the works it cites.
Quip: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and De Sa, C · 2023
Closest in time.
RedPajama: An open source recipe to reproduce llama training dataset, 2023
Computer, T · 2023
Closest in time.
SparseGPT: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Closest in time.
GPTQ code, 2023
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Closest in time.
MegaBlocks: Efficient sparse training with mixture-of-experts
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dettmers, T. and Zettlemoyer, L · 2022
Cited alongside, same era.
LLM.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Cited alongside, same era.
GLaM: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
GPTQ: Accurate post-training compression for generative pretrained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Koishekenov, Y., Nikoulina, V., and Berard, A · 2022
Cited alongside, same era.
The Optimal BERT Surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Cited alongside, same era.
Gale, T., Narayanan, D., Young, C., and Zaharia, M · 2023
Closest in time.
T5x, 2023
Google · 2023
Closest in time.
Tutel: Adaptive mixture-of-experts at scale
Hwang, C., Cui, W., Xiong, Y., Yang, Z., Liu, Z., Hu, H., Wang, Z., Salas, R., Jose, J., Ram, P., et al · 2023
Closest in time.
Finequant: Unlocking efficiency with fine-grained weight-only quantization for llms
Kim, Y. J., Henry, R., Fahim, R., and Awadalla, H. H · 2023
Closest in time.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Closest in time.
Efficient GPU kernels for n:m-sparse weights in deep learning
Lin, B., Zheng, N., Wang, L., Cao, S., Ma, L., Zhang, Q., Zhu, Y., Cao, T., Xue, J., Yang, Y., et al · 2023
Closest in time.
From sparse to soft mixtures of experts
Puigcerver, J., Riquelme, C., Mustafa, B., and Houlsby, N · 2023
Closest in time.
ZeroQuant-FP: A leap forward in llms post-training w4a8 quantization using floating-point formats
Wu, X., Yao, Z., and He, Y · 2023
Closest in time.
Edgemoe: Fast on-device inference of moe-based large language models
Yi, R., Guo, L., Wei, S., Zhou, A., Wang, S., and Xu, M · 2023
Closest in time.
Boost transformer-based language models with gpu-friendly sparsity and quantization
Yu, C., Chen, T., and Gan, Z · 2023
Closest in time.