Fetching the paper…
Reading the bibliography…
Transformers, driven by attention mechanisms, form the foundation of large language models (LLMs).
Fast transformer decoding: One write-head is all you need
Shazeer, N · 1911
Earlier work this paper cites.
Principles of runtime support for parallel processors
Mirchandaney, R., Saltz, J. H., Smith, R. M., Nicol, D. M., and Crowley, K · 1988
Earlier work this paper cites.
The preprocessed doacross loop
Saltz, J. H. and Mirchandaney, R · 1991
Earlier work this paper cites.
Runtime compilation techniques for data partitioning and communication schedule reuse
Ponnusamy, R., Saltz, J. H., and Choudhary, A. N · 1993
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Sparsity: Optimization framework for sparse matrix kernels
Im, E., Yelick, K. A., and Vuduc, R. W · 2004
Earlier work this paper cites.
Parallel sparse matrix-vector and matrix-transpose-vector multiplication using compressed sparse blocks
Buluç, A., Fineman, J. T., Frigo, M., Gilbert, J. R., and Leiserson, C. E · 2009
Earlier work this paper cites.
Spgrid: a sparse paged grid structure applied to adaptive smoke simulation
Setaluri, R., Aanjaneya, M., Bauer, S., and Sifakis, E · 2014
Earlier work this paper cites.
Gpu kernels for block-sparse weights
Gray, S., Radford, A., and Kingma, D. P · 2017
Earlier work this paper cites.
Block-sparse recurrent neural networks
Narang, S., Undersander, E., and Diamos, G. F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Online normalizer calculation for softmax
Milakov, M. and Gimelshein, N · 2018
Earlier work this paper cites.
Ragged tensors | tensorflow core
Tensorflow Developers · 2018
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H. T., and Cox, D · 2019
Earlier work this paper cites.
Efficient tensor core-based GPU kernels for structured sparsity under reduced precision
Chen, Z., Qu, Z., Liu, L., Ding, Y., and Xie, Y · 2021
Earlier work this paper cites.
FasterTransformer
NVIDIA · 2021
Earlier work this paper cites.
Fusedmm: A unified sddmm-spmm kernel for graph embedding and graph neural networks
Rahman, M. K., Sujon, M. H., and Azad, A · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Efficient quantized sparse matrix operations on tensor cores
Li, S., Osawa, K., and Hoefler, T · 2022
Earlier work this paper cites.
Micikevicius, P., Stosic, D., Burgess, N., Cornea, M., Dubey, P., Grisenthwaite, R., Ha, S., Heinecke, A., Judd, P., Kamalu, J., Mellempudi, N., Oberman, S. F., Shoeybi, M., Siu, M. Y., and Wu, H · 2022
Earlier work this paper cites.
Sequential aggregation and rematerialization: Distributed full-batch training of graph neural networks on large graphs
Mostafa, H · 2022
Earlier work this paper cites.
Nvidia hopper architecture in-depth, 2022
NVIDIA · 2022
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2022
Earlier work this paper cites.
Orca: A distributed serving system for transformer-based generative models
Yu, G., Jeong, J. S., Kim, G., Kim, S., and Chun, B · 2022
Earlier work this paper cites.
Understanding GNN computational graph: A coordinated computation, io, and memory perspective
Zhang, H., Yu, Z., Dai, G., Huang, G., Ding, Y., Xie, Y., and Wang, Y · 2022
Earlier work this paper cites.
GQA: training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Cited alongside, same era.
Flash-decoding for long-context inference, 2023
Dao, T., Haziza, D., Massa, F., and Sizov, G · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Mixed-input matrix multiplication performance optimizations
Gupta, M · 2024
Later among the works it cites.
Flexattention: The flexibility of pytorch with the performance of flashattention, Aug 2024
He, H., Guessous, D., Liang, Y., and Dong, J · 2024
Later among the works it cites.
Flashdecoding++: Faster large language model inference with asynchronization, flat gemm optimization, and heuristics
Hong, K., Dai, G., Xu, J., Mao, Q., Li, X., Liu, J., chen, k., Dong, Y., and Wang, Y · 2024
Later among the works it cites.
Hydragen: High-throughput LLM inference with shared prefixes
Juravsky, J., Brown, B. C. A., Ehrlich, R. S., Fu, D. Y., Ré, C., and Mirhoseini, A · 2024
Later among the works it cites.
Pod-attention: Unlocking full prefill-decode overlap for faster llm inference
Kamath, A. K., Prabhu, R., Mohan, J., Peter, S., Ramjee, R., and Panwar, A · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lai, R., Shao, J., Feng, S., Lyubomirsky, S. S., Hou, B., Lin, W., Ye, Z., Jin, H., Jin, Y., Liu, J., Jin, L., Cai, Y., Jiang, Z., Wu, Y., Park, S., Srivastava, P., Roesch, J. G., Mowry, T. C., and Chen, T · 2023
Cited alongside, same era.
Blockwise parallel transformers for large context models
Liu, H. and Abbeel, P · 2023
Cited alongside, same era.
Ring attention with blockwise transformers for near-infinite context
Liu, H., Zaharia, M., and Abbeel, P · 2023
Cited alongside, same era.
Stream-k: Work-centric parallel decomposition for dense matrix-matrix multiplication on the GPU
Osama, M., Merrill, D., Cecka, C., Garland, M., and Owens, J. D · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M. H. M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2023
Cited alongside, same era.
CUTLASS, January 2023
Thakkar, V., Ramani, P., Cecka, C., Shivam, A., Lu, H., Yan, E., Kosaian, J., Hoemmen, M., Wu, H., Kerr, A., Nicely, M., Merrill, D., Blasig, D., Qiao, F., Majcher, P., Springer, P., Hohnerbach, M., Wang, J., and Gupta, M · 2023
Cited alongside, same era.
TC-GNN: bridging sparse GNN computation and dense tensor cores on gpus
Wang, Y., Feng, B., Wang, Z., Huang, G., and Ding, Y · 2023
Cited alongside, same era.
Later among the works it cites.
Parrot: Efficient serving of LLM-based applications with semantic variable
Lin, C., Han, Z., Zhang, C., Yang, Y., Yang, F., Chen, C., and Qiu, L · 2024
Later among the works it cites.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Zhang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., Shi, C., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2024
Later among the works it cites.
Accelerating PyTorch with CUDA Graphs
Nguyen, V., Carilli, M., Eryilmaz, S. B., Singh, V., Lin, M., Gimelshein, N., Desmaison, A., and Yang, E · 2024
Later among the works it cites.
Nvdsl: Simplifying tensor cores with python-driven mlir metaprogramming
Ozen, G · 2024
Later among the works it cites.
vattention: Dynamic memory management for serving llms without pagedattention, 2024
Prabhu, R., Nayak, A., Mohan, J., Ramjee, R., and Panwar, A · 2024
Later among the works it cites.
attention-gym
PyTorch-Labs · 2024
Later among the works it cites.
Theory, analysis, and best practices for sigmoid self-attention
Ramapuram, J., Danieli, F., Dhekane, E. G., Weers, F., Busbridge, D., Ablin, P., Likhomanenko, T., Digani, J., Gu, Z., Shidani, A., and Webb, R · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Rivière, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., Ferret, J., Liu, P., Tafti, P., Friesen, A., Casbon, M., Ramos, S., Kumar, R., Lan, C. L., Jerome, S., Tsitsulin, A., Vieillard, N., Stanczyk, P., Girgin, S., Momchev, N., Hoffman, M., Thakoor, S., Grill, J., Neyshabur, B., Bachem, O., Walton, A., Severyn, A., Parrish, A., Ahmad, A., Hutchison, A., Abdagic, A., Carl, A., Shen, A., Brock, A., Coenen, A., Laforge, A., Paterson, A., Bastian, B., Piot, B., Wu, B., Royal, B., Chen, C., Kumar, C., Perry, C., Welty, C., Choquette-Choo, C. A., Sinopalnikov, D., Weinberger, D., Vijaykumar, D., Rogozinska, D., Herbison, D., Bandy, E., Wang, E., Noland, E., Moreira, E., Senter, E., Eltyshev, E., Visin, F., Rasskin, G., Wei, G., Cameron, G., Martins, G., Hashemi, H., Klimczak-Plucinska, H., Batra, H., Dhand, H., Nardini, I., Mein, J., Zhou, J., Svensson, J., Stanway, J., Chan, J., Zhou, J. P., Carrasqueira, J., Iljazi, J., Becker, J., Fernandez, J., van Amersfoort, J., Gordon, J., Lipschultz, J., Newlan, J., Ji, J., Mohamed, K., Badola, K., Black, K., Millican, K., McDonell, K., Nguyen, K., Sodhia, K., Greene, K., Sjösund, L. L., Usui, L., Sifre, L., Heuermann, L., Lago, L., and McNealus, L · 2024
Later among the works it cites.
Lean attention: Hardware-aware scalable attention mechanism for the decode-phase of transformers
Sanovar, R., Bharadwaj, S., Amant, R. S., Rühle, V., and Rajmohan, S · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Later among the works it cites.
ThunderKittens: A Simple Embedded DSL for AI kernels, May 2024
Spector, B., Singhal, A., Arora, S., and Re, C · 2024
Later among the works it cites.
QUEST: query-aware sparsity for efficient long-context LLM inference
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Later among the works it cites.
A multi-level superoptimizer for tensor programs
Wu, M., Cheng, X., Padon, O., and Jia, Z · 2024
Later among the works it cites.
Open Release of Grok-1
xAI · 2024
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training, 2024
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2024
Later among the works it cites.
Chunkattention: Efficient self-attention with prefix-aware KV cache and two-phase partition
Ye, L., Tao, Z., Huang, Y., and Li, Y · 2024
Later among the works it cites.
Relayattention for efficient large language model serving with long system prompts
Zhu, L., Wang, X., Zhang, W., and Lau, R. W. H · 2024
Later among the works it cites.
Optimizing and characterizing high-throughput low-latency LLM inference in MLCEngine, Oct 2024
MLC Community · 2026
Closest in time.
NVIDIA TensorRT-LLM, 2023a
NVIDIA · 2026
Closest in time.