Fetching the paper…
Reading the bibliography…
The prohibitive sizes of Large Language Models (LLMs) today make it difficult to deploy them on memory-constrained edge devices.
Dither Signals and Their Effect on Quantization Noise
L. Schuchman · 1964
Earlier work this paper cites.
Dithered quantizers
R. Gray and T. Stockham · 1993
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge, 2009
L. Bentivogli, I. Dagan, H. T. Dang, D. Giampiccolo, and B. Magnini · 2009
Earlier work this paper cites.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
N. Halko, P. G. Martinsson, and J. A. Tropp · 2011
Earlier work this paper cites.
The winograd schema challenge
H. J. Levesque, E. Davis, and L. Morgenstern · 2012
Earlier work this paper cites.
Optimal exact least squares rank minimization
S. Xiang, Y. Zhu, X. Shen, and J. Ye · 2012
Earlier work this paper cites.
Matrix Computations
G. H. Golub and C. F. van Loan · 2013
Earlier work this paper cites.
Matrix inversion using cholesky decomposition
A. Krishnamoorthy and D. Menon · 2013
Earlier work this paper cites.
Pointer Sentinel Mixture Models, 2016
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
The sphere packing problem in dimension 8 8
M. Viazovska · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
R. Vershynin · 2018
Earlier work this paper cites.
WinoGrande: An Adversarial Winograd Schema Challenge at Scale, 2019
S. Keisuke, L. B. Ronan, B. Chandra, and C. Yejin · 2019
Earlier work this paper cites.
Why are big data matrices approximately low rank?
M. Udell and A. Townsend · 2019
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, 2019
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2019
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi · 2020
Cited alongside, same era.
Up or Down? Adaptive Rounding for Post-Training Quantization
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models, 2020
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus, 2021
J. Dodge, M. Sap, A. Marasović, W. Agnew, G. Ilharco, D. Groeneveld, M. Mitchell, and M. Gardner · 2021
Cited alongside, same era.
Matrix Compression via Randomized Low Rank and Low Precision Factorization
R. Saha, V. Srivastava, and M. Pilanci · 2023
Later among the works it cites.
The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction, 2023
P. Sharma, J. T. Ash, and D. Misra · 2023
Later among the works it cites.
Redpajama: an open dataset for training large language models, October 2023
Together Computer · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Gugger, L. Debut, T. Wolf, P. Schmid, Z. Mueller, S. Mangrulkar, M. Sun, and B. Bossan · 2022
Cited alongside, same era.
Language model compression with weighted low-rank factorization
Y.-C. Hsu, T. Hua, S. Chang, Q. Lou, Y. Shen, and H. Jin · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
J. Chee, Y. Cai, V. Kuleshov, and C. D. Sa · 2023
Cited alongside, same era.
QLoRA: Efficient Finetuning of Quantized LLMs
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer · 2023
Cited alongside, same era.
OPTQ: Accurate Quantization for Generative Pre-trained Transformers
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2023
Cited alongside, same era.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2023
Later among the works it cites.
ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation, 2023
Z. Yao, X. Wu, C. Li, S. Youn, and Y. He · 2023
Later among the works it cites.
ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models, 2023
Z. Yuan, Y. Shang, Y. Song, Q. Wu, Y. Yan, and G. Sun · 2023
Later among the works it cites.
Extreme Compression of Large Language Models via Additive Quantization, 2024
V. Egiazarian, A. Panferov, D. Kuznedelev, E. Frantar, A. Babenko, and D. Alistarh · 2024
Closest in time.
How Good Are Low-bit Quantized LLaMA3 Models? An Empirical Study, 2024
W. Huang, X. Ma, H. Qin, X. Zheng, C. Lv, H. Chen, J. Luo, X. Qi, X. Liu, and M. Magno · 2024
Closest in time.
Dora: Weight-decomposed low-rank adaptation, 2024
S.-Y. Liu, C.-Y. Wang, H. Yin, P. Molchanov, Y.-C. F. Wang, K.-T. Cheng, and M.-H. Chen · 2024
Closest in time.
The era of 1-bit llms: All large language models are in 1.58 bits, 2024
S. Ma, H. Wang, L. Ma, L. Wang, W. Wang, S. Huang, L. Dong, R. Wang, J. Xue, and F. Wei · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta AI · 2024
Closest in time.
Building a gpt-style llm classifier from scratch, 2024
S. Raschka · 2024
Closest in time.
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, 2024
A. Tseng, J. Chee, Q. Sun, V. Kuleshov, and C. D. Sa · 2024
Closest in time.
LQER: Low-Rank Quantization Error Reconstruction for LLMs, 2024
C. Zhang, J. Cheng, G. A. Constantinides, and Y. Zhao · 2024
Closest in time.