Fetching the paper…
Reading the bibliography…
Optimizing deep learning algorithms currently requires slow, manual derivation, potentially leaving much performance untapped.
Fast transformer decoding: One write-head is all you need, 2019
Noam Shazeer · 1911
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, et al · 1912
Earlier work this paper cites.
The input/output complexity of sorting and related problems
Alok Aggarwal and Jeffrey Vitter, S · 1988
Earlier work this paper cites.
The deep learning compiler: A comprehensive survey, 2020
Mingzhen Li, Yi Liu, Xiaoyan Liu, et al · 2002
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2004
Earlier work this paper cites.
Data movement is all you need: A case study on optimizing transformers, 2021
Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, et al · 2007
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, et al · 2017
Earlier work this paper cites.
What your DRAM power models are not telling you: Lessons from a detailed experimental study, 2018
Saugata Ghose, Abdullah Giray Yağlıkçı, Raghav Gupta, et al · 2018
Earlier work this paper cites.
Backprop as functor: A compositional perspective on supervised learning
Brendan Fong, David I. Spivak, and Rémy Tuyéras · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Philippe Tillet, H. T. Kung, and David Cox · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
NVIDIA a100 tensor core GPU architecture overview, 2020
NVIDIA · 2020
Earlier work this paper cites.
Xla: Compiling machine learning for peak performance
Amit Sabne · 2020
Earlier work this paper cites.
Categorical foundations of gradient-based learning, 2021
G. S. H. Cruttwell, Bruno Gavranović, Neil Ghani, et al · 2021
Earlier work this paper cites.
A survey of quantization methods for efficient neural network inference, 2021
Amir Gholami, Sehoon Kim, Zhen Dong, et al · 2021
Cited alongside, same era.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness, 2022
Tri Dao, Daniel Y. Fu, Stefano Ermon, et al · 2022
Cited alongside, same era.
NVIDIA h100 tensor core GPU architecture overview, 2022
NVIDIA · 2022
Cited alongside, same era.
Markov categories and entropy, 2022
Paolo Perrone · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, et al · 2022
An introduction to string diagrams for computer scientists, 2023
Robin Piedeleu and Fabio Zanasi · 2023
Later among the works it cites.
SDXL: Improving latent diffusion models for high-resolution image synthesis, 2023
Dustin Podell, Zion English, Kyle Lacey, et al · 2023
Later among the works it cites.
Category-theoretic data structures and algorithms for learning polynomial circuits, 2023
Paul William Wilson · 2023
Later among the works it cites.
Co-design of complex systems: From autonomy to future mobility systems, 2023
Gioele Zardini · 2023
Later among the works it cites.
The llama 3 herd of models, 2024
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robust diagrams for deep learning architectures: Applications and theory, 2023
Vincent Abbott · 2023
Cited alongside, same era.
GQA: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, et al · 2023
Cited alongside, same era.
AMD CDNA 3 architecture, 2023
AMD · 2023
Cited alongside, same era.
Developing CUDA kernels for accelerated matrix multiplication on NVIDIA hopper architecture using the CUTLASS library
Ganesh Bikshandi and Jay Shah · 2023
Cited alongside, same era.
A case study in CUDA kernel fusion: Implementing FlashAttention-2 on NVIDIA hopper architecture using the CUTLASS library
Ganesh Bikshandi, Jay Shah, and Colfax Research · 2023
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Cited alongside, same era.
GPTQ: Accurate post-training quantization for generative pre-trained transformers, 2023
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2023
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, et al · 2024
Closest in time.
Amir Gholami, Zhewei Yao, Sehoon Kim, et al · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, et al · 2024
Closest in time.
Benchmarking and dissecting the nvidia hopper GPU architecture, 2024
Weile Luo, Ruibo Fan, Zeyu Li, et al · 2024
Closest in time.
PTX ISA 8.5, 2024
NVIDIA · 2024
Closest in time.
A note on the algebra of CuTe layouts
Jay Shah · 2024
Closest in time.
FlashAttention-3: Fast and accurate attention with asynchrony and low-precision, 2024
Jay Shah, Ganesh Bikshandi, Ying Zhang, et al · 2024
Closest in time.
NVIDIA blackwell architecture technical overview, 2025
NVIDIA · 2025
Closest in time.