Fetching the paper…
Reading the bibliography…
Quantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning.
Fixup initialization: Residual learning without normalization
Zhang, H · 1901
Earlier work this paper cites.
Lewis, M · 1910
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding
Nie, Y · 1910
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Bar-Haim, R · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I · 2007
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Giampiccolo, D · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L · 2009
Earlier work this paper cites.
Binarybert: Pushing the limit of bert quantization
Bai, H · 2012
Earlier work this paper cites.
The winograd schema challenge
Levesque, H · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R · 2013
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S · 2016
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P · 2016
Cited alongside, same era.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, D · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A · 2017
Cited alongside, same era.
Narayan, S · 2018
Q-bert: Hessian based ultra low precision quantization of bert
Shen, S · 2020
Later among the works it cites.
Training verifiers to solve math word problems
Cobbe, K · 2021
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Hu, E. J · 2021
Later among the works it cites.
Towards efficient post-training quantization of pre-trained language models
Bai, H · 2022
Later among the works it cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A · 2019
Cited alongside, same era.
Neural network acceptability judgments
Warstadt, A · 2019
Cited alongside, same era.
Q8bert: Quantized 8bit bert
Zafrir, O · 2019
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M · 2020
Cited alongside, same era.
Frantar, E · 2022
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T · 2023
Closest in time.
Lmflow: An extensible toolkit for finetuning and inference of large foundation models
Diao, S · 2023
Closest in time.
Losparse: Structured compression of large language models based on low-rank and sparse approximation
Li, Y · 2023
Closest in time.
Llm-qat: Data-free quantization aware training for large language models
Liu, Z · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G · 2023
Closest in time.
Adaptive budget allocation for parameter-efficient fine-tuning
Zhang, Q · 2023
Closest in time.