Fetching the paper…
Reading the bibliography…
In recent years, Transformer-based language models have become the standard approach for natural language processing tasks.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2018
Earlier work this paper cites.
Glow: Graph lowering compiler techniques for neural networks
Rotem, N., Fix, J., Abdulrasool, S., Catron, G., Deng, S., Dzhabarov, R., Gibson, N., Hegeman, J., Lele, M., Levenstein, R., et al · 2018
Earlier work this paper cites.
Efficient 8-bit quantization of transformer neural machine language translation model
Bhandare, A., Sripathi, V., Karkada, D., Menon, V., Choi, S., Datta, K., and Saletore, V · 2019
Earlier work this paper cites.
Optimizing dnn computation with relaxed graph substitutions
Jia, Z., Thomas, J., Warszawski, T., Gao, M., Zaharia, M., and Aiken, A · 2019
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Distilling task-specific knowledge from bert into simple neural networks
Tang, R., Lu, Y., Liu, L., Mou, L., Vechtomova, O., and Lin, J · 2019
Cited alongside, same era.
Q8bert: Quantized 8bit bert
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Compressing bert: Studying the effects of weight pruning on transfer learning
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M · 2021
Later among the works it cites.
Channel permutations for n: m sparsity
Pool, J. and Yu, C · 2021
Later among the works it cites.
Accelerating inference with sparsity using the nvidia ampere architecture and nvidia tensorrt
Pool, J., Sawarkar, A., and Rodge, J · 2021
Later among the works it cites.
Sparsednn: Fast sparse deep learning inference on cpus
Wang, Z · 2021
Later among the works it cites.
Prune once for all: Sparse pre-trained language models
Zafrir, O., Larey, A., Boudoukh, G., Shen, H., and Wasserblat, M · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gordon, M. A., Duh, K., and Andrews, N · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Cited alongside, same era.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D · 2020
Cited alongside, same era.
Sparsert: Accelerating unstructured sparsity on gpus for deep learning inference
Wang, Z · 2020
Cited alongside, same era.
I-bert: Integer-only bert quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Cited alongside, same era.
Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H · 2021
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Later among the works it cites.
Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al · 2022
Later among the works it cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Later among the works it cites.