Fetching the paper…
Reading the bibliography…
The attention layer, a core component of Transformer-based LLMs, brings out inefficiencies in current GPU systems due to its low operational intensity and the substantial memory requirements of KV caches.
Long short-term memory
Alex Graves · 2012
Earlier work this paper cites.
Recurrent neural networks
Stephen Grossberg · 2013
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Efficient attention: Attention with linear complexities
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li · 2021
Earlier work this paper cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Earlier work this paper cites.
Unleashing the potential of pim: Accelerating large batched inference of transformer-based generative models
Jaewan Choi, Jaehyun Park, Kwanhee Kyung, Nam Sung Kim, and Jung Ho Ahn · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
https://openai.com/index/gpt-4/, Accessed: 2024-11-09
Gpt4 · 2024
Cited alongside, same era.
https://www.llama.com/, Accessed: 2024-11-09
Llama3.2 · 2024
Cited alongside, same era.
Attacc! unleashing the power of pim for batched transformer-based generative model inference
Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn · 2024
Cited alongside, same era.
Fastdecode: High-throughput gpu-efficient llm serving using heterogeneous pipelines
Jiaao He and Jidong Zhai · 2024
Cited alongside, same era.
https://www.nvidia.com/en-us/data-center/a100/, Accessed: 2024-10-21
Nvidia a100 · 2024
Cited alongside, same era.
https://www.nvidia.com/en-us/data-center/h100/, Accessed: 2024-10-21
Sungmin Yun, Kwanhee Kyung, Juhwan Cho, Jaewan Choi, Jongmin Kim, Byeongho Kim, Sukhan Lee, Kyomin Sohn, and Jung Ho Ahn · 2024
Later among the works it cites.
https://developer.nvidia.com/nsight-systems, Accessed: 2024-10-29
Nvidia nsight systems · 2024
Later among the works it cites.
https://developer.nvidia.com/nsight-compute, Accessed: 2024-10-29
Nvidia nsight compute · 2024
Later among the works it cites.
https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices, Accessed: 2024-10-29
Llm inference performance engineering: Best practices · 2024
Later among the works it cites.
https://news.skhynix.com/sk-hynix-begins-volume-production-of-the-world-first-12-layer-hbm3e/, Accessed: 2024-10-29
Sk hynix begins volume production of the world’s first 12-layer hbm3e · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nvidia h100 · 2024
Cited alongside, same era.
Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing
Guseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi, Sanghyeon Lee, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan, and Jongse Park · 2024
Cited alongside, same era.
Palm: scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sashank Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel
Cited in the paper.
https://github.com/mosaicml/llm-foundry/tree/main/scripts/train/benchmarking, Accessed: 2024-10-29
Mpt training benchmarks · 2024
Later among the works it cites.
https://www.semianalysis.com/p/100000-h100-clusters-power-network, Accessed: 2024-10-29
100,000 h100 clusters: Power, network topology, ethernet vs infiniband, reliability, failures, checkpointing · 2024
Later among the works it cites.