Fetching the paper…
Reading the bibliography…
The demand for inference on extremely large scale LLMs has seen enormous growth in the recent months.
“Pointer sentinel mixture models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2016
Earlier work this paper cites.
“Attention is all you need”
A. Vaswani et al · 2017
Earlier work this paper cites.
P. Micikevicius et al · 2017
Earlier work this paper cites.
“Benchmarking TPU, GPU, and CPU platforms for deep learning”
Y.. Wang, G.-Y. Wei and D. Brooks · 2019
Earlier work this paper cites.
“Q8BERT: Quantized 8bit BERT”
O. Zafrir, G. Boudoukh, P. Izsak and M. Wasserblat · 2019
Earlier work this paper cites.
“GOBO: quantizing attention-based NLP models for low latency and energy efficient inference”
A.. Zadeh, I. Edo, O.. Awad and A. Moshovos · 2020
Earlier work this paper cites.
“Q-BERT: hessian based ultra low precision quantization of BERT”
S. Shen et al · 2020
Earlier work this paper cites.
“TernaryBERT: distillation-aware ultra-low bit BERT”
W. Zhang et al · 2020
Earlier work this paper cites.
“Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point”
Bita Darvish et al · 2020
Earlier work this paper cites.
“Bottleneck transformers for visual recognition”
A. Srinivas et al · 2021
Earlier work this paper cites.
“OPT: open pre-trained transformer language models”
Susan Zhang et al · 2022
Cited alongside, same era.
“Block Floating Point (BFP) for Efficient Deep Neural Net Inference”
I. Lyubomirsky and X. Wang · 2022
Cited alongside, same era.
“Block Format Error Bounds and Optimal Block Size Selection”
I. Soloveychik, I. Lyubomirsky, X. Wang and S. Bhoja · 2022
Cited alongside, same era.
“FP8 formats for deep learning”
Paulius Micikevicius et al · 2022
Cited alongside, same era.
“LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale”
Tim Dettmers, Mike Lewis, Younes Belkada and Luke Zettlemoyer · 2022
“Microscaling Data Formats for Deep Learning”
Bita Rouhani et al · 2023
Later among the works it cites.
“FP8-lm: Training FP8 large language models”
Houwen Peng et al · 2023
Later among the works it cites.
OpenAI · 2024
Closest in time.
Albert. Jiang et al · 2024
Closest in time.
“Gemma: open Models Based on Gemini Research and Technology”
Gemma Team et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Llama 2: open Foundation and Fine-Tuned Chat Models”
Hugo Touvron et al · 2023
Cited alongside, same era.
“GPTQ: accurate Post-Training Quantization for Generative Pre-trained Transformers”
Elias Frantar, Saleh Ashkboos, Torsten Hoefler and Dan Alistarh · 2023
Cited alongside, same era.
“SmoothQuant: accurate and Efficient Post-Training Quantization for Large Language Models”
Guangxuan Xiao et al · 2023
Cited alongside, same era.
“MX Pytorch Emulation Library”
Microsoft · 2023
Cited alongside, same era.
“KVQuant: towards 10 Million Context Length LLM Inference with KV Cache Quantization”
Coleman Hooper et al · 2024
Closest in time.
“LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens”
Yiran Ding et al · 2024
Closest in time.
“The Era of 1-bit LLMs: all large language lodels are in 1.58 bits”
Shuming Ma et al · 2024
Closest in time.
“RoFormer: enhanced transformer with rotary position embedding”
Jianlin Su et al · 2024
Closest in time.