Fetching the paper…
Reading the bibliography…
The size and compute characteristics of modern large language models have led to an increased interest in developing specialized kernels tailored for particular training and inference workloads.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax, 2018
Milakov, M. and Gimelshein, N · 2018
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H. T., and Cox, D · 2019
Earlier work this paper cites.
Xla : Compiling machine learning for peak performance, 2020
Sabne, A · 2020
Earlier work this paper cites.
Glu variants improve transformer
Shazeer, N · 2020
Earlier work this paper cites.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Earlier work this paper cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
GPTQ: Accurate post-training compression for generative pretrained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Fast inference from transformers via speculative decoding, 2022
Leviathan, Y., Kalman, M., and Matias, Y · 2022
Earlier work this paper cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., De Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Cited alongside, same era.
Quip: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and De Sa, C. M · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Cited alongside, same era.
Flash-decoding for long-context inference, 2023
Dao, T., Haziza, D., Massa, F., and Sizov, G · 2023
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Later among the works it cites.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Later among the works it cites.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
DeepSeek-AI · 2024
Later among the works it cites.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al · 2024
Later among the works it cites.
Eagle: Speculative sampling requires rethinking feature uncertainty, 2024
Li, Y., Wei, F., Zhang, C., and Zhang, H · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pre-rmsnorm and pre-crmsnorm transformers: equivalent and efficient pre-ln transformers
Jiang, Z., Gu, J., Zhu, H., and Pan, D · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Hydra: Sequentially-dependent draft heads for medusa decoding, 2024
Ankner, Z., Parthasarathy, R., Nrusimha, A., Rinard, C., Ragan-Kelley, J., and Brandon, W · 2024
Cited alongside, same era.
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation
Ansel, J., Yang, E., He, H., Gimelshein, N., Jain, A., Voznesensky, M., Bao, B., Bell, P., Berard, D., Burovski, E., Chauhan, G., Chourdia, A., Constable, W., Desmaison, A., DeVito, Z., Ellison, E., Feng, W., Gong, J., Gschwind, M., Hirsh, B., Huang, S., Kalambarkar, K., Kirsch, L., Lazos, M., Lezcano, M., Liang, Y., Liang, J., Lu, Y., Luk, C., Maher, B., Pan, Y., Puhrsch, C., Reso, M., Saroufim, M., Siraichi, M. Y., Suk, H., Suo, M., Tillet, P., Wang, E., Wang, X., Wen, W., Zhang, S., Zhao, X., Zhou, K., Zou, R., Mathews, A., Chanan, G., Wu, P., and Chintala, S · 2024
Cited alongside, same era.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024a
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T
Cited in the paper.
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T
Cited in the paper.
Flexattention: The flexibility of pytorch with the performance of flashattention, 8 2024
Guessous, D., Liang, Y., Dong, J., and He, H · 2025
Closest in time.
Thundermla: Flashmla, faster and fused-er!, 2023
Spector, B., Singhal, A., Fu, D., and Ré, C · 2025
Closest in time.
Mirage: A multi-level superoptimizer for tensor programs
Wu, M., Cheng, X., Liu, S., Shi, C., Ji, J., Ao, K., Velliengiri, P., Miao, X., Padon, O., and Jia, Z · 2025
Closest in time.