Fetching the paper…
Reading the bibliography…
This paper presents MoE-Infinity, an efficient MoE inference system designed for personal machines with limited GPU memory capacity.
On the closest packing of spheres in n dimensions
Rankin, R. A · 1947
Earlier work this paper cites.
Covering spheres with spheres
Dumer, I · 2007
Earlier work this paper cites.
SwapAdvisor: Pushing deep learning beyond the GPU memory limit via smart swapping
Huang, C., Jin, G., and Li, J · 2020
Earlier work this paper cites.
Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
Fedus, W., Zoph, B., and Shazeer, N · 2021
Earlier work this paper cites.
Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning
Ren, J., Luo, J., Wu, K., Zhang, M., Jeon, H., and Li, D · 2021
Earlier work this paper cites.
DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., and He, Y · 2022
Earlier work this paper cites.
No language left behind: Scaling human-centered machine translation, 2022
Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Gonzalez, G. M., Hansanti, P., Hoffman, J., Jarrett, S., Sadagopan, K. R., Rowe, D., Spruit, S., Tran, C., Andrews, P., Ayan, N. F., Bhosale, S., Edunov, S., Fan, A., Gao, C., Goswami, V., Guzmán, F., Koehn, P., Mourachko, A., Ropers, C., Saleem, S., Schwenk, H., and Wang, J · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
Orca: A distributed serving system for transformer-based generative models
Yu, G., Jeong, J. S., Kim, G., Kim, S., and Chun, B · 2022
Cited alongside, same era.
Optimizing dynamic neural networks with brainstorm
Cui, W., Han, Z., Ouyang, L., Wang, Y., Zheng, N., Ma, L., Yang, Y., Yang, F., Xue, J., Qiu, L., Zhou, L., Chen, Q., Tan, H., and Guo, M · 2023
Cited alongside, same era.
Fast inference of mixture-of-experts language models with offloading, 2023
Eliseev, A. and Mazur, D · 2023
Cited alongside, same era.
Fast and efficient model serving using multi-gpus with direct-host-access
Jeong, J., Baek, S., and Ahn, J · 2023
Cited alongside, same era.
DeepUM: Tensor migration and prefetching in unified memory
Jung, J., Kim, J., and Lee, J · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
The faiss library
Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazaré, P.-E., Lomeli, M., Hosseini, L., and Jégou, H · 2024
Closest in time.
DeepSpeed-FastGen: High-throughput text generation for LLMs via MII and DeepSpeed-Inference, 2024
Holmes, C., Tanaka, M., Wyatt, M., Awan, A. A., Rasley, J., Rajbhandari, S., Aminabadi, R. Y., Qin, H., Bakhtiari, A., Kurilenko, L., and He, Y · 2024
Closest in time.
Text generation inference
HuggingFace · 2024
Closest in time.
Mixtral of experts, 2024
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
Infinigen: Efficient generative inference of large language models with dynamic KV cache management
Lee, W., Lee, J., Seo, J., and Sim, J · 2024
Closest in time.
TensorRT-LLM
NVIDIA · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single GPU
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Cited alongside, same era.
DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model, 2024
DeepSeek-AI · 2024
Cited alongside, same era.
Snowflake Arctic: The best LLM for enterprise AI — efficiently intelligent, truly open
Snowflake AI Research
Cited in the paper.
Introducing DBRX: A new state-of-the-art open LLM
The Mosaic Research Team
Cited in the paper.
Closest in time.
Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters”, February 2024
Team, Q · 2024
Closest in time.
Open release of Grok-1
XAI · 2024
Closest in time.