Fetching the paper…
Reading the bibliography…
Improving the deployment efficiency of transformer-based language models has been challenging given their high computation and memory cost.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcinkiewicz, M. A · 1994
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Recognizing textual entailment: Models and applications
Dagan, I., Roth, D., Sammons, M., and Zanzotto, F. M · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2017
Earlier work this paper cites.
First quora dataset release: Question pairs.(2017)
Iyer, S., Dandekar, N., and Csernai, K · 2017
Earlier work this paper cites.
Exploring the regularity of sparse structure in convolutional neural networks
Mao, H., Han, S., Pool, J., Li, W., Liu, X., Wang, Y., and Dally, W. J · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
CUTLASS: Fast Linear Algebra in CUDA C++
NVIDIA · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R · 2017
Earlier work this paper cites.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł · 2018
Earlier work this paper cites.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary!: topic-aware convolutional neural networks for extreme summarization
Narayan, S., Martins, A., Sordoni, A., Bachman, P., Courville, A., and Bengio, Y · 2018
Earlier work this paper cites.
Model compression via distillation and quantization
Polino, A., Pascanu, R., and Alistarh, D · 2018
Earlier work this paper cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2018
Cited alongside, same era.
HAWQ: Hessian aware quantization of neural networks with mixed-precision
Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2019
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2019
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
A survey of quantization methods for efficient neural network inference
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2021
Later among the works it cites.
I-bert: Integer-only bert quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Later among the works it cites.
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M · 2021
Later among the works it cites.
Post-training quantization for vision transformer
Liu, Z., Wang, Y., Han, K., Zhang, W., Ma, S., and Gao, W · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Cited alongside, same era.
Bert rediscovers the classical nlp pipeline
Tenney, I., Das, D., and Pavlick, E · 2019
Cited alongside, same era.
Binarybert: Pushing the limit of bert quantization
Bai, H., Zhang, W., Hou, L., Shang, L., Jin, J., Jiang, X., Liu, Q., Lyu, M., and King, I · 2020
Cited alongside, same era.
Extremely low bit transformer quantization for on-device neural machine translation
Chung, I., Kim, B., Choi, Y., Kwon, S. J., Jeon, Y., Park, B., Kim, S., and Lee, D · 2020
Cited alongside, same era.
Mishra, A., Latorre, J. A., Pool, J., Stosic, D., Stosic, D., Venkatesh, G., Yu, C., and Micikevicius, P · 2021
Later among the works it cites.
Employing CUDA Graphs in a Dynamic Environment
NVIDIA · 2021
Later among the works it cites.
LEAP: Learnable Pruning for Transformer-based Models
Yao, Z., Wu, X., Ma, L., Shen, S., Keutzer, K., Mahoney, M. W., and He, Y · 2021
Later among the works it cites.
Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Zhang, M., Awan, A. A., Li, C., Li, D., Zheng, E., Rasley, J., Smith, S., Ruwase, O., et al · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2022
Later among the works it cites.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Frantar, E. and Alistarh, D · 2022
Later among the works it cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Later among the works it cites.
Compressing pre-trained transformers via low-bit nxm sparsity for natural language understanding
Holmes, C., Zhang, M., He, Y., and Wu, B · 2022
Later among the works it cites.
Dq-bart: Efficient sequence-to-sequence model via joint distillation and quantization
Li, Z., Wang, Z., Tan, M., Nallapati, R., Bhatia, P., Arnold, A., Xiang, B., and Roth, D · 2022
Later among the works it cites.
Mkq-bert: Quantized bert with 4-bits weights and activations
Tang, H., Zhang, X., Liu, K., Zhu, J., and Kang, Z · 2022
Later among the works it cites.
Extreme compression for pre-trained transformers made simple and efficient
Wu, X., Yao, Z., Zhang, M., Li, C., and He, Y · 2022
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Later among the works it cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Later among the works it cites.
GPU workstation for deep learning
Lambda · 2023
Closest in time.
FasterTransformer
NVIDIA · 2023
Closest in time.
A comprehensive study on post-training quantization for large language models
Yao, Z., Li, C., Wu, X., Youn, S., and He, Y · 2023
Closest in time.