Fetching the paper…
Reading the bibliography…
We introduce Mirage, the first multi-level superoptimizer for tensor programs.
Probabilistic algorithms for sparse polynomials
Richard Zippel · 1979
Earlier work this paper cites.
Fast probabilistic algorithms for verification of polynomial identities
J. T. Schwartz · 1980
Earlier work this paper cites.
Superoptimizer: a look at the smallest program
Henry Massalin · 1987
Earlier work this paper cites.
Automatic generation of peephole superoptimizers
Sorav Bansal and Alex Aiken · 2006
Earlier work this paper cites.
Z3: An efficient smt solver
Leonardo De Moura and Nikolaj Bjørner · 2008
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Halide: A language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe · 2013
Earlier work this paper cites.
Stochastic superoptimization
Eric Schkufza, Rahul Sharma, and Alex Aiken · 2013
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer · 2014
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
https://developer.nvidia.com/cublas , 2016
Dense Linear Algebra on GPUs · 2016
Earlier work this paper cites.
Automatically scheduling halide image processing pipelines
Ravi Teja Mullapudi, Andrew Adams, Dillon Sharlet, Jonathan Ragan-Kelley, and Kayvon Fatahalian · 2016
Earlier work this paper cites.
https://www.tensorflow.org/xla , 2017
Xla: Optimizing compiler for tensorflow · 2017
Earlier work this paper cites.
https://pytorch.org , 2017
Tensors and Dynamic neural networks in Python with strong GPU acceleration · 2017
Earlier work this paper cites.
https://developer.nvidia.com/tensorrt , 2017
NVIDIA TensorRT: Programmable inference accelerator · 2017
Earlier work this paper cites.
TVM: end-to-end optimization stack for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Haichen Shen, Eddie Q. Yan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Learning to optimize tensor programs
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Nvidia tensor core programmability, performance & precision
Stefano Markidis, Steven Wei Der Chien, Erwin Laure, Ivy Bo Peng, and Jeffrey S. Vetter · 2018
Earlier work this paper cites.
https://github.com/NVIDIA/cutlass , 2019
Nvidia/cutlass: Cuda templates for linear algebra subroutines · 2019
Cited alongside, same era.
https://www.tensorflow.org/guide/graph_optimization , 2019
Tensorflow graph optimization with grappler · 2019
Cited alongside, same era.
Taso: Optimizing deep learning computation with automatic generation of graph substitutions
Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken · 2019
Cited alongside, same era.
Beyond data and model parallelism for deep neural networks
Zhihao Jia, Matei Zaharia, and Alex Aiken · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
https://huggingface.co/Laurie/llama7b-lora-merged/tree/main , 2023
Llama-7b-lora · 2023
Later among the works it cites.
https://www.nvidia.com/en-us/data-center/h100/ , 2023
Nvidia h100 tensor core gpu · 2023
Later among the works it cites.
https://triton-lang.org/main/getting-started/tutorials/06-fused-attention.html , 2023
A Triton implementation of the FlashAttention2 algorithm · 2023
Later among the works it cites.
Falcon-40B: an open large language model with state-of-the-art performance
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo · 2023
Later among the works it cites.
Flash-decoding for long-context inference, 2023
Tri Dao, Daniel Haziza, Francisco Massa, and Grigory Sizov · 2023
Later among the works it cites.
Graphene: An ir for optimized tensor computations on gpus
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Philippe Tillet, H. T. Kung, and David Cox · 2019
Cited alongside, same era.
Root mean square layer normalization, 2019
Biao Zhang and Rico Sennrich · 2019
Cited alongside, same era.
https://github.com/NVIDIA/FasterTransformer , 2020
Transformer related optimizations · 2020
Cited alongside, same era.
Ansor : Generating high-performance tensor programs for deep learning
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica · 2020
Cited alongside, same era.
Flextensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system
Size Zheng, Yun Liang, Shuo Wang, Renze Chen, and Kaiwen Sheng · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
PET: Optimizing tensor programs with partially equivalent transformations and automated corrections
Haojie Wang, Jidong Zhai, Mingyu Gao, Zixuan Ma, Shizhi Tang, Liyan Zheng, Yuanzhi Li, Kaiyuan Rong, Yuanyong Chen, and Zhihao Jia · 2021
Cited alongside, same era.
Bastian Hagedorn, Bin Fan, Hanfeng Chen, Cris Cecka, Michael Garland, and Vinod Grover · 2023
Later among the works it cites.
Aspen: Breaking operator barriers for efficient parallelization of deep neural networks
Jongseok Park, Kyungmin Bin, Gibum Park, Sangtae Ha, and Kyunghan Lee · 2023
Later among the works it cites.
Welder: Scheduling deep learning memory access via tile-graph
Yining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo, Fan Yang, and Lidong Zhou · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
EINNET: Optimizing tensor programs with Derivation-Based transformations
Liyan Zheng, Haojie Wang, Jidong Zhai, Muyan Hu, Zixuan Ma, Tuowei Wang, Shuhong Huang, Xupeng Miao, Shizhi Tang, Kezhao Huang, and Zhihao Jia · 2023
Later among the works it cites.
Flashdecoding++: Faster large language model inference on gpus, 2024
Ke Hong, Guohao Dai, Jiaming Xu, Qiuli Mao, Xiuhong Li, Jun Liu, Kangdi Chen, Yuhan Dong, and Yu Wang · 2024
Closest in time.
Optimal kernel orchestration for tensor programs with korch
Muyan Hu, Ashwin Venkatram, Shreyashri Biswas, Balamurugan Marimuthu, Bohan Hou, Gabriele Oliaro, Haojie Wang, Liyan Zheng, Xupeng Miao, Jidong Zhai, and Zhihao Jia · 2024
Closest in time.
nGPT: Normalized transformer with representation learning on the hypersphere, 2024
Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, and Boris Ginsburg · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models, 2024
Chameleon Team · 2024
Closest in time.
The llama 3 herd of models, 2024
The Llama 3 team · 2024
Closest in time.
Graphpipe: Improving performance and scalability of dnn training with graph pipeline parallelism
Byungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim, Sunghyun Park, Neeraj Aggarwal, Colin Unger, Daiyaan Arfeen, Peiyuan Liao, Xupeng Miao, Mohammad Alizadeh, Gregory R. Ganger, Tianqi Chen, and Zhihao Jia · 2025
Closest in time.
Identity testing for circuits with exponentiation gates, 2025
Jiatu Li and Mengdi Wu · 2025
Closest in time.