Fetching the paper…
Reading the bibliography…
Transformers are the state-of-the-art architecture for large language models, and a key to their scalability is the strategic usage of low-precision arithmetic.
Root mean square layer normalization
B. Zhang and R. Sennrich · 1910
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
N. Shazeer · 1911
Earlier work this paper cites.
A theory of condition
J. R. Rice · 1966
Earlier work this paper cites.
Norms on direct sums and tensor products
P. Lancaster and H. K. Farahat · 1972
Earlier work this paper cites.
Mixed, componentwise, and structured condition numbers
I. Gohberg and I. Koltracht · 1993
Earlier work this paper cites.
Topics in Matrix Analysis
R. A. Horn and C. R. Johnson · 1994
Earlier work this paper cites.
Applied Functional Analysis: Main Principles and their Applications
E. Zeidler · 1995
Earlier work this paper cites.
Accuracy and Stability of Numerical Algorithms
N. J. Higham · 2002
Earlier work this paper cites.
GLU variants improve transformer
N. Shazeer · 2002
Earlier work this paper cites.
On layer normalization in the transformer architecture
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu · 2002
Earlier work this paper cites.
ReZero is all you need: Fast convergence at large depth
T. Bachlechner, B. P. Majumder, H. Mao, G. Cottrell, and J. McAuley · 2003
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2004
Earlier work this paper cites.
Up or down? adaptive rounding for post-training quantization
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort · 2004
Earlier work this paper cites.
The Lipschitz constant of self-attention
H. Kim, G. Papamakarios, and A. Mnih · 2006
Earlier work this paper cites.
Functions of Matrices: Theory and Computation
N. J. Higham · 2008
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy · 2010
Earlier work this paper cites.
Matrix Analysis
R. A. Horn and C. R. Johnson · 2012
Earlier work this paper cites.
Layer normalization
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
D. Lin, S. Talathi, and S. Annapureddy · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky · 2017
Earlier work this paper cites.
Parseval networks: Improving robustness to adversarial examples
M. Cisse, P. Bojanowski, E. Grave, Y. Dauphin, and N. Usunier · 2017
Earlier work this paper cites.
Analytical guarantees on numerical precision of deep neural networks
C. Sakr, Y. Kim, and N. Shanbhag · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Efficient processing of deep neural networks: A tutorial and survey
V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko · 2018
Earlier work this paper cites.
NVIDIA tensor core programmability, performance & precision
S. Markidis, S. W. Der Chien, E. Laure, I. B. Peng, and J. S. Vetter · 2018
Earlier work this paper cites.
Mixed precision training
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, et al · 2018
Earlier work this paper cites.
Handbook of Floating-Point Arithmetic
J.-M. Muller, N. Brisebarre, F. De Dinechin, C.-P. Jeannerod, V. Lefevre, G. Melquiond, N. Revol, D. Stehlé, S. Torres, et al · 2018
Earlier work this paper cites.
Evaluating the robustness of neural networks: An extreme value theory approach
T.-W. Weng, H. Zhang, P.-Y. Chen, J. Yi, D. Su, Y. Gao, C.-J. Hsieh, and L. Daniel · 2018
Cited alongside, same era.
Sorting out Lipschitz function approximation
C. Anil, J. Lucas, and R. Grosse · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. Le, and R. Salakhutdinov · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Matrix Differential Calculus with Applications in Statistics and Econometrics
J. R. Magnus and H. Neudecker · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
RoFormer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2023
Later among the works it cites.
SmoothQuant: Accurate and efficient post-training quantization for large language models
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2023
Later among the works it cites.
Adversarial ink: Componentwise backward error attacks on deep learning
L. Beerens and D. J. Higham · 2024
Later among the works it cites.
How smooth is attention?
V. Castin, P. Ablin, and G. Peyré · 2024
Later among the works it cites.
Accumulator-aware post-training quantization for large language models
I. Colbert, F. Grob, G. Franco, J. Zhang, and R. Saab · 2024
Later among the works it cites.
Bounds on nonlinear errors for variance computation with stochastic rounding
E.-M. El Arar, D. Sohier, P. de Oliveira Castro, and E. Petit · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robustness may be at odds with accuracy
D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry · 2019
Cited alongside, same era.
Mixed precision block fused multiply-add: Error analysis and application to GPU tensor cores
P. Blanchard, N. J. Higham, F. Lopez, T. Mary, and S. Pranesh · 2020
Cited alongside, same era.
Adversarial attacks via backward error analysis
T. Beuzeville, P. Boudier, A. Buttari, S. Gratton, T. Mary, and S. Pralet · 2021
Cited alongside, same era.
Accurately computing the log-sum-exp and softmax functions
P. Blanchard, D. J. Higham, and N. J. Higham · 2021
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Y. Dong, J.-B. Cordonnier, and A. Loukas · 2021
Cited alongside, same era.
Regularisation of neural networks by enforcing Lipschitz continuity
H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree · 2021
Cited alongside, same era.
Later among the works it cites.
The Llama 3 herd of models
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al · 2024
Later among the works it cites.
Mixtral of experts
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Later among the works it cites.
AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, and S. Han · 2024
Later among the works it cites.
Deepseek-v3 Technical report
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al · 2024
Later among the works it cites.
Matrix derivatives: Why and where did it go wrong?
J. R. Magnus · 2024
Later among the works it cites.
Falcon2-11b technical report
Q. Malartic, N. R. Chowdhury, R. Cojocaru, M. Farooq, G. Campesan, Y. A. D. Djilali, S. Narayan, A. Singh, M. Velikanov, B. E. A. Boussaha, et al · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al · 2024
Later among the works it cites.
FlashAttention-3: Fast and accurate attention with asynchrony and low-precision
J. Shah, G. Bikshandi, Y. Zhang, V. Thakkar, P. Ramani, and T. Dao · 2024
Later among the works it cites.
DeepNet: Scaling transformers to 1,000 layers
H. Wang, S. Ma, L. Dong, S. Huang, D. Zhang, and F. Wei · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis · 2024
Later among the works it cites.
Qwen2.5 Technical report
A. Yang, B. Yang, B. Zhang, B. Hui, et al · 2024
Later among the works it cites.
Analysis of floating-point matrix multiplication computed via integer arithmetic
A. Abdelfattah, J. Dongarra, M. Fasi, M. Mikaitis, and F. Tisseur · 2025
Closest in time.
Quantization error propagation: Revisiting layer-wise post-training quantization
Y. Arai and Y. Ichikawa · 2025
Closest in time.
Mixed precision accumulation for neural network inference guided by componentwise forward error analysis
E.-M. El Arar, S.-I. Filip, T. Mary, and E. Riccietti · 2025
Closest in time.
HIGGS: Pushing the limits of large language model quantization via the linearity theorem
V. Malinovskii, A. Panferov, I. Ilin, H. Guo, P. Richtárik, and D. Alistarh · 2025
Closest in time.
Training transformers with enforced Lipschitz constants
L. Newhouse, R. P. Hess, F. Cesista, A. Zahorodnii, J. Bernstein, and P. Isola · 2025
Closest in time.
Defeating the training-inference mismatch via FP16
P. Qi, Z. Liu, X. Zhou, T. Pang, C. Du, W. S. Lee, and M. Lin · 2025
Closest in time.
2 OLMo 2 furious
P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y. Gu, S. Huang, M. Jordan, et al · 2025
Closest in time.
Pay attention to attention distribution: A new local Lipschitz bound for transformers
N. Yudin, A. Gaponov, S. Kudriashov, and M. Rakhuba · 2025
Closest in time.
Deterministic and probabilistic rounding error analysis of neural networks in floating-point arithmetic
T. Beuzeville, A. Buttari, S. Gratton, and T. Mary · 2026
Closest in time.
LAMP: Look-ahead mixed-precision inference of large language models
S. Budzinskiy, M. Gloser, T. Yilmaz, Y. H. Tham, Y. Lin, W. Fang, F. Wu, and P. Petersen · 2026
Closest in time.
Exact attention sensitivity and the geometry of transformer stability
S. M. Emadi · 2026
Closest in time.
Ministral 3
A. H. Liu, K. Khandelwal, S. Subramanian, V. Jouault, A. Rastogi, A. Sadé, A. Jeffares, A. Jiang, A. Cahill, A. Gavaudan, et al · 2026
Closest in time.