Fetching the paper…
Reading the bibliography…
eXmY is a novel data type for quantization of ML models.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
C. Clark, K. Lee, M. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 1905
Earlier work this paper cites.
HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision, 2019b
Z. Dong, Z. Yao, A. Gholami, M. Mahoney, and K. Keutzer · 1905
Earlier work this paper cites.
Deep Learning Recommendation Model for Personalization and Recommendation Systems, 2019
M. Naumov, D. Mudigere, et al · 1906
Earlier work this paper cites.
HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks, 2019a
Z. Dong, Z. Yao, Y. Cai, D. Arfeen, A. Gholami, M. W. Mahoney, and K. Keutzer · 1911
Earlier work this paper cites.
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
J. H. Clark, E. Choi, M. Collins, D. Garrette, T. Kwiatkowski, V. Nikolaev, and J. Palomaki · 2003
Earlier work this paper cites.
Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation, 2020
H. Wu, P. Judd, X. Zhang, M. Isaev, and P. Micikevicius · 2004
Earlier work this paper cites.
Language Models are Few-Shot Learners, 2020
T. B. Brown, B. Mann, et al · 2005
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
M. Roemmele, C. Bejan, and A. Gordon · 2011
Earlier work this paper cites.
HAWQV3: Dyadic Neural Network Quantization, 2021
Z. Yao, Z. Dong, Z. Zheng, A. Gholami, J. Yu, E. Tan, L. Wang, Q. Huang, Y. Wang, M. W. Mahoney, and K. Keutzer · 2011
Earlier work this paper cites.
The Winograd Schema Challenge
H. J. Levesque, E. Davis, and L. Morgenstern · 2012
Earlier work this paper cites.
Semantic Parsing on Freebase from Question-Answer Pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories
N. Mostafazadeh, N. Chambers, et al · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, et al · 2016
Earlier work this paper cites.
Beating Floating Point at its Own Game: Posit Arithmetic, 2017
J. L. Gustafson and I. Yonemoto · 2017
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding Comprehension Dataset From Examinations
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, et al · 2017
Earlier work this paper cites.
QuAC : Question Answering in Context
E. Choi, H. He, M. Iyyer, M. Yatskar, W. Yih, Y. Choi, P. Liang, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Cited alongside, same era.
ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension
S. Zhang, X. Liu, J. Liu, J. Gao, K. Duh, and B. V. Durme · 2018
Cited alongside, same era.
DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
D. Dua, Y. Wang, et al · 2019
A survey of quantization methods for efficient neural network inference
A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer · 2022
Later among the works it cites.
PERCIVAL: Open-Source Posit RISC-V Core With Quire Capability
D. Mallasén, R. Murillo, et al · 2022
Later among the works it cites.
FP8 Formats for Deep Learning, 2022
P. Micikevicius, D. Stosic, N. Burgess, M. Cornea, P. Dubey, R. Grisenthwaite, S. Ha, A. Heinecke, P. Judd, J. Kamalu, N. Mellempudi, S. Oberman, M. Shoeybi, M. Siu, and H. Wu · 2022
Later among the works it cites.
Optimal clipping and magnitude-aware differentiation for improved quantization-aware training
C. Sakr, S. Dai, R. Venkatesan, B. Zimmer, W. Dally, and B. Khailany · 2022
Later among the works it cites.
ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers, 2022
Z. Yao, R. Y. Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BFloat16: The secret to high performance on Cloud TPUs, 2019
Google · 2019
Cited alongside, same era.
Natural Questions: A Benchmark for Question Answering Research
T. Kwiatkowski, J. Palomaki, et al · 2019
Cited alongside, same era.
Mixed precision training with 8-bit floating point
N. Mellempudi, S. Srinivasan, D. Das, and B. Kaul · 2019
Cited alongside, same era.
IEEE Standard for Floating-Point Arithmetic
Microprocessor Standards Committee · 2019
Cited alongside, same era.
WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations, 2019
M. T. Pilehvar and J. Camacho-Collados · 2019
Cited alongside, same era.
HellaSwag: Can a Machine Really Finish Your Sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
PIQA: Reasoning about Physical Commonsense in Natural Language
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi · 2020
Cited alongside, same era.
Later among the works it cites.
R. Anil, A. M. Dai, et al · 2023
Later among the works it cites.
Unit scaling: Out-of-the-box low-precision training
C. Blake, D. Orr, and C. Luschi · 2023
Later among the works it cites.
QLoRA: Efficient Finetuning of Quantized LLMs, 2023
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer · 2023
Later among the works it cites.
OPTQ: Accurate Quantization for Generative Pre-trained Transformers
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2023
Later among the works it cites.
N. P. Jouppi, G. Kurian, et al · 2023
Later among the works it cites.
Llm-fp4: 4-bit floating-point quantized transformers
S.-y. Liu, Z. Liu, X. Huang, P. Dong, and K.-T. Cheng · 2023
Later among the works it cites.
OCP Microscaling Formats (MX) Specification, Sep 2023
B. D. Rouhani, N. Garegrat, et al · 2023
Later among the works it cites.
Augmenting hessians with inter-layer dependencies for mixed-precision post-training quantization
C. J. Schaefer, N. Lambert-Shirzad, X. Zhang, C. Chou, T. Jablin, J. Li, E. Guo, C. Stanton, S. Joshi, and Y. E. Wang · 2023
Later among the works it cites.
X. Wei, Y. Zhang, Y. Li, X. Zhang, R. Gong, J. Guo, and X. Liu · 2023
Later among the works it cites.
Quarot: Outlier-free 4-bit inference in rotated llms
S. Ashkboos, A. Mohtashami, M. L. Croci, B. Li, M. Jaggi, D. Alistarh, T. Hoefler, and J. Hensman · 2024
Closest in time.
Quantizable transformers: Removing outliers by helping attention heads do nothing
Y. Bondarenko, M. Nagel, and T. Blankevoort · 2024
Closest in time.
FP8 Quantization: The Power of the Exponent, 2024
A. Kuzmin, M. V. Baalen, Y. Ren, M. Nagel, J. Peters, and T. Blankevoort · 2024
Closest in time.
LLaMA 3, 2024
Meta · 2024
Closest in time.
NVIDIA Blackwell Architecture, 2024
NVIDIA · 2024
Closest in time.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, 2024
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2024
Closest in time.