Fetching the paper…
Reading the bibliography…
The deployment of Large Language Models (LLMs) on edge devices is increasingly important to enhance on-device intelligence.
Llvm: A compilation framework for lifelong program analysis & transformation
Chris Lattner and Vikram Adve · 2004
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
Tvm: an automated end-to-end optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Learning to optimize tensor programs
Tianqi Chen, Lianmin Zheng, Eddie Q. Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
High throughput matrix-matrix multiplication between asymmetric bit-width operands, 2020
Dibakar Gope, Jesse Beu, and Matthew Mattina · 2020
Earlier work this paper cites.
Multiplying matrices without multiplying
Davis Blalock and John Guttag · 2021
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Cited alongside, same era.
The case for 4-bit precision: k-bit inference scaling laws
Tim Dettmers and Luke Zettlemoyer · 2023
Cited alongside, same era.
Deepgemm: Accelerated ultra low-precision inference on cpu architectures using lookup tables
Darshan C Ganji, Saad Ashfaq, Ehsan Saboori, Sudhakar Sah, Saptarshi Mitra, Mohammadhossein Askarihemmat, Alexander Hoffman, Ahmed Hassanien, and Mathieu Leonardon · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Awq: Activation-aware weight quantization for llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han · 2023
Bitnet: Scaling 1-bit transformers for large language models
Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Huaijie Wang, Lingxiao Ma, Fan Yang, Ruiping Wang, Yi Wu, and Furu Wei · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Later among the works it cites.
https://blogs.microsoft.com/blog/2024/05/20/introducing-copilot-pcs/#_ftn2
Introducing Copilot+ PCs · 2024
Closest in time.
Quip: 2-bit quantization of large language models with guarantees, 2024
Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher De Sa · 2024
Closest in time.
Bitdistiller: Unleashing the potential of sub-4-bit llms via self-distillation, 2024
Dayou Du, Yijia Zhang, Shijie Cao, Jiaqi Guo, Ting Cao, Xiaowen Chu, and Ningyi Xu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Look-up mai gemm: Increasing ai gemms performance by nearly 2.5 x via msgemm
Saeed Maleki · 2023
Cited alongside, same era.
Lut-gemm: Quantized matrix multiplication based on luts for efficient inference in large-scale generative language models, 2023
Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee · 2023
Cited alongside, same era.
Lut-nn: Empower efficient neural network inference with centroid learning and table lookup
Xiaohu Tang, Yang Wang, Ting Cao, Li Lyna Zhang, Qi Chen, Deng Cai, Yunxin Liu, and Mao Yang · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Cited alongside, same era.
https://github.com/dmlc/dlpack
DLPack
Cited in the paper.
https://github.com/intel/neural-compressor
Intel Neural Compressor
Cited in the paper.
https://huggingface.co/TheBloke/Llama-2-7B-GGUF
Llama-2-7B GGUF Models
Cited in the paper.
Marlin: a fast 4-bit inference kernel for medium batchsizes
Elias Frantar and Dan Alistarh · 2024
Closest in time.
Onebit: Towards extremely low-bit large language models
Yuzhuang Xu, Xu Han, Zonghan Yang, Shuo Wang, Qingfu Zhu, Zhiyuan Liu, Weidong Liu, and Wanxiang Che · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al · 2024
Closest in time.