Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) models face deployment challenges due to their large parameter counts and computational demands.
Bounds on multiprocessing timing anomalies
Graham, R · 1966
Earlier work this paper cites.
Roofline: an insightful visual performance model for multicore architectures
Williams, S., Waterman, A., and Patterson, D · 2009
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Low-bit quantization of neural networks for efficient inference
Choukroun, Y., Kravchik, E., Yang, F., and Kisilev, P · 2019
Earlier work this paper cites.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2019
Earlier work this paper cites.
Haq: Hardware-aware automated quantization with mixed precision
Wang, K., Liu, Z., Lin, Y., Lin, J., and Han, S · 2019
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Earlier work this paper cites.
Squeezellm: Dense-and-sparse quantization
Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
CUTLASS, January 2023
Thakkar, V., Ramani, P., Cecka, C., Shivam, A., Lu, H., Yan, E., Kosaian, J., Hoemmen, M., Wu, H., Kerr, A., Nicely, M., Merrill, D., Blasig, D., Qiao, F., Majcher, P., Springer, P., Hohnerbach, M., Wang, J., and Gupta, M · 2023
Cited alongside, same era.
Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x
Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Wang, Z., Shen, L., Wang, A., Li, Y., Su, T., Yang, Z., and Tang, J · 2023
Cited alongside, same era.
Quarot: Outlier-free 4-bit inference in rotated llms
Ashkboos, S., Mohtashami, A., Croci, M. L., Li, B., Cameron, P., Jaggi, M., Alistarh, D., Hoefler, T., and Hensman, J · 2024
Cited alongside, same era.
Skvq: Sliding-window key and value cache quantization for large language models
Examining post-training quantization for mixture-of-experts: A benchmark
Li, P., Jin, X., Cheng, Y., and Chen, T · 2024
Later among the works it cites.
Olmoe: Open mixture-of-experts language models
Muennighoff, N., Soldaini, L., Groeneveld, D., Lo, K., Morrison, J., Min, S., Shi, W., Walsh, P., Tafjord, O., Lambert, N., et al · 2024
Later among the works it cites.
Massive activations in large language models
Sun, M., Chen, X., Kolter, J. Z., and Liu, Z · 2024
Later among the works it cites.
Hobbit: A mixed precision expert offloading system for fast moe inference
Tang, P., Liu, J., Hou, X., Pu, Y., Wang, J., Heng, P.-A., Li, C., and Guo, M · 2024
Later among the works it cites.
Introducing DBRX: A New State-of-the-Art Open LLM, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Duanmu, H., Yuan, Z., Li, X., Duan, J., Zhang, X., and Lin, D · 2024
Cited alongside, same era.
Marlin: Mixed-precision auto-regressive parallel inference on large language models
Frantar, E., Castro, R. L., Chen, J., Hoefler, T., and Alistarh, D · 2024
Cited alongside, same era.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
Cited alongside, same era.
Mc-moe: Mixture compressor for mixture-of-experts llms gains more
Huang, W., Liao, Y., Liu, J., He, R., Tan, H., Zhang, S., Li, H., Liu, S., and Qi, X · 2024
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Cited alongside, same era.
Who says elephants can’t run: Bringing large scale moe models into cloud scale production
Kim, Y. J., Henry, R., Fahim, R., and Awadalla, H. H
Cited in the paper.
Who says elephants can’t run: Bringing large scale moe models into cloud scale production
Kim, Y. J., Henry, R., Fahim, R., and Awadalla, H. H
Cited in the paper.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S
Cited in the paper.
The Mosaic Research Team · 2024
Later among the works it cites.
Openmoe: An early effort on open mixture-of-experts language models
Xue, F., Zheng, Z., Fu, Y., Ni, J., Zheng, Z., Zhou, W., and You, Y · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Wkvquant: Quantizing weight and key/value cache for large language models gains more
Yue, Y., Yuan, Z., Duanmu, H., Zhou, S., Wu, J., and Nie, L · 2024
Later among the works it cites.
Atom: Low-bit quantization for efficient and accurate llm serving
Zhao, Y., Lin, C.-Y., Zhu, K., Ye, Z., Chen, L., Zheng, S., Ceze, L., Krishnamurthy, A., Chen, T., and Kasikci, B · 2024
Later among the works it cites.