Exploring the limits of transfer learning with a unified text-to-text transformer
Original
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
Turing-NLG: A 17-billion-parameter language model by microsoft
Rosset, C · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Original
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Later among the works it cites.
Megatron-LM: Training multi-billion parameter language models using gpu model parallelism
Original
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Later among the works it cites.
Patient knowledge distillation for bert model compression
Original
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Original
Tan, M. and Le, Q. V · 2019
Later among the works it cites.
Distilling task-specific knowledge from bert into simple neural networks
Original
Tang, R., Lu, Y., Liu, L., Mou, L., Vechtomova, O., and Lin, J · 2019
Later among the works it cites.
Well-read students learn better: On the importance of pre-training compact models
Original
Turc, I., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
HAQ: Hardware-aware automated quantization
Wang, K., Liu, Z., Lin, Y., Lin, J., and Han, S · 2019
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Later among the works it cites.
Q8BERT: Quantized 8bit bert
Original
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Later among the works it cites.
Cortex-M, https://developer.arm.com/ip-products/processors/cortex-m, 2020
ARM · 2020
Later among the works it cites.
Binarybert: Pushing the limit of bert quantization
Original
Bai, H., Zhang, W., Hou, L., Shang, L., Jin, J., Jiang, X., Liu, Q., Lyu, M., and King, I · 2020
Later among the works it cites.
Language models are few-shot learners
Original
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Original
Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., and Smith, N · 2020
Later among the works it cites.
Training with quantization noise for extreme fixed-point compression
Original
Fan, A., Stock, P., Graham, B., Grave, E., Gribonval, R., Jegou, H., and Joulin, A · 2020
Later among the works it cites.
Compressing bert: Studying the effects of weight pruning on transfer learning
Original
Gordon, M. A., Duh, K., and Andrews, N · 2020
Later among the works it cites.
GShard: Scaling giant models with conditional computation and automatic sharding
Original
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Later among the works it cites.
Ladabert: Lightweight adaptation of bert through hybrid model compression
Original
Mao, Y., Wang, Y., Wu, C., Zhang, C., Wang, Y., Yang, Y., Zhang, Q., Tong, Y., and Bai, J · 2020
Later among the works it cites.
Fixed encoder self-attention patterns in transformer-based machine translation
Original
Raganato, A., Scherrer, Y., and Tiedemann, J · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Original
Sanh, V., Wolf, T., and Rush, A. M · 2020
Later among the works it cites.
Q-BERT: Hessian based ultra low precision quantization of bert
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Later among the works it cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Original
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D · 2020
Later among the works it cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Original
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., and Zhou, M · 2020
Later among the works it cites.
Bert-of-theseus: Compressing bert by progressive module replacing
Original
Xu, C., Zhou, W., Ge, T., Wei, F., and Zhou, M · 2020
Later among the works it cites.
HAWQV3: Dyadic neural network quantization
Original
Yao, Z., Dong, Z., Zheng, Z., Gholami, A., Yu, J., Tan, E., Wang, L., Huang, Q., Wang, Y., Mahoney, M. W., and Keutzer, K · 2020
Later among the works it cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Zadeh, A. H., Edo, I., Awad, O. M., and Moshovos, A · 2020
Later among the works it cites.
Ternarybert: Distillation-aware ultra-low bit bert
Original
Zhang, W., Hou, L., Yin, Y., Shang, L., Chen, X., Jiang, X., and Liu, Q · 2020
Later among the works it cites.
Kdlsq-bert: A quantized bert combining knowledge distillation with learned step size quantization
Original
Jin, J., Liang, C., Wu, T., Zou, L., and Gan, Z · 2021
Closest in time.
https://github.com/kssteven418/i-bert, 2021
Kim, S · 2021
Closest in time.