BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., and Kang, J · 2020
Later among the works it cites.
Task-specific objectives of pre-trained language models for dialogue adaptation
Original
Li, J., Zhang, Z., Zhao, H., Zhou, X., and Zhou, X · 2020
Later among the works it cites.
FastBERT: a self-distilling BERT with adaptive inference time
Liu, W., Zhou, P., Wang, Z., Zhao, Z., Deng, H., and Ju, Q · 2020
Later among the works it cites.
Q-BERT: Hessian based ultra low precision quantization of BERT
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D · 2020
Later among the works it cites.
Structured pruning of large language models
Wang, Z., Wohlwend, J., and Lei, T · 2020
Later among the works it cites.
BERT-of-theseus: Compressing BERT by progressive module replacing
Xu, C., Zhou, W., Ge, T., Wei, F., and Zhou, M · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
You, Y., Li, J., Reddi, S. J., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C · 2020
Later among the works it cites.
Active learning with sampling by uncertainty and density for data annotations
Zhu, J., Wang, H., Tsou, B. K., and Ma, M · 2020
Later among the works it cites.
EarlyBERT: Efficient BERT training via early-bird lottery tickets
Chen, X., Cheng, Y., Wang, S., Gan, Z., Wang, Z., and Liu, J · 2021
Closest in time.
DeBERTa: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W · 2021
Closest in time.
I-BERT: integer-only BERT quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Closest in time.
Primer: Searching for efficient transformers for language modeling, 2021
So, D. R., Mańke, W., Liu, H., Dai, Z., Shazeer, N., and Le, Q. V · 2021
Closest in time.
Scale efficiently: Insights from pre-training and fine-tuning transformers, 2021
Tay, Y., Dehghani, M., Rao, J., Fedus, W., Abnar, S., Chung, H. W., Narang, S., Yogatama, D., Vaswani, A., and Metzler, D · 2021
Closest in time.
Finetuned language models are zero-shot learners
Original
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Closest in time.
Deep co-training with task decomposition for semi-supervised domain adaptation
Yang, L., Wang, Y., Gao, M., Shrivastava, A., Weinberger, K. Q., Chao, W.-L., and Lim, S.-N · 2021
Closest in time.
Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections
Zhong, R., Lee, K., Zhang, Z., and Klein, D · 2021
Closest in time.