Fetching the paper…
Reading the bibliography…
Computational complexity and overthinking problems have become the bottlenecks for pre-training language models (PLMs) with millions or even trillions of parameters.
“The ATIS spoken language systems pilot corpus,”
Hemphill et al., · 1990
Earlier work this paper cites.
“Stack overflow creative commons data dump,”
Jeff Atwood, · 2009
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey E. Hinton et al., · 2015
Earlier work this paper cites.
“Branchynet: Fast inference via early exiting from deep neural networks,”
Teerapittayanon et al., · 2016
Earlier work this paper cites.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin et al., · 2018
Earlier work this paper cites.
Mostafa Dehghani et al., · 2018
Earlier work this paper cites.
“How to stop off-the-shelf deep neural networks from overthinking,”
Yigitcan Kaya et al., · 2018
Earlier work this paper cites.
“GLUE: A multi-task benchmark and analysis platform for natural language understanding,”
Alex Wang et al., · 2018
Earlier work this paper cites.
Alice Coucke et al., · 2018
Earlier work this paper cites.
“SGM: sequence generation model for multi-label classification,”
Pengcheng Yang et al., · 2018
Earlier work this paper cites.
“Panlp at mediqa 2019: Pre-trained language models, transfer learning and knowledge distillation,”
Wei Zhu, Xiaofeng Zhou, Keqiang Wang, Xun Luo, Xiepeng Li, Yuan Ni, and Guo Tong Xie, · 2019
Earlier work this paper cites.
“ALBERT: A lite BERT for self-supervised learning of language representations,”
Zhenzhong Lan et al., · 2019
Earlier work this paper cites.
“Xlnet: Generalized autoregressive pretraining for language understanding,”
Zhilin Yang et al., · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized BERT pretraining approach,”
Yinhan Liu et al., · 2019
Earlier work this paper cites.
“Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter,”
Victor Sanh et al., · 2019
Cited alongside, same era.
“Reducing transformer depth on demand with structured dropout,”
Angela Fan et al., · 2019
Cited alongside, same era.
“Are sixteen heads really better than one?,”
Paul Michel et al., · 2019
Cited alongside, same era.
“Patient knowledge distillation for BERT model compression,”
Siqi Sun et al., · 2019
Cited alongside, same era.
“Decoupled weight decay regularization,”
Ilya Loshchilov et al., · 2019
Cited alongside, same era.
“Global attention decoder for chinese spelling error correction,”
Zhao Guo, Yuan Ni, Keqiang Wang, Wei Zhu, and Guo Tong Xie, · 2021
Later among the works it cites.
“MVP-BERT: Multi-vocab pre-training for Chinese BERT,”
Wei Zhu, · 2021
Later among the works it cites.
“I-BERT: integer-only BERT quantization,”
Sehoon Kim et al., · 2021
Later among the works it cites.
“Autonlu: Architecture search for sentence and cross-sentence attention modeling with re-designed search space,”
Wei Zhu, · 2021
Later among the works it cites.
“Automatic student network search for knowledge distillation,”
Zhexi Zhang, Wei Zhu, Junchi Yan, Peng Gao, and Guowang Xie, · 2021
Later among the works it cites.
“Discovering better model architectures for medical query understanding,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Mvp-bert: Redesigning vocabularies for chinese bert and multi-vocab pretraining,”
Wei Zhu, · 2020
Cited alongside, same era.
“BERT loses patience: Fast and robust inference with early exit,”
Wangchunshu Zhou et al., · 2020
Cited alongside, same era.
“Bert-of-theseus: Compressing bert by progressive module replacing,”
Canwen et al., · 2020
Cited alongside, same era.
“Fastbert: a self-distilling BERT with adaptive inference time,”
Weijie Liu et al., · 2020
Cited alongside, same era.
“Deebert: Dynamic early exiting for accelerating BERT inference,”
Ji Xin et al., · 2020
Cited alongside, same era.
“The right tool for the job: Matching model and instance complexities,”
Roy Schwartz et al., · 2020
Cited alongside, same era.
“Binarybert: Pushing the limit of bert quantization,”
Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jing Jin, Xin Jiang, Qun Liu, Michael R. Lyu, and Irwin King, · 2020
Cited alongside, same era.
Wei Zhu, Yuan Ni, Xiaoling Wang, and Guo Tong Xie, · 2021
Later among the works it cites.
“BERxiT: Early exiting for BERT with better fine-tuning and extension to regression,”
Ji et al. Xin, · 2021
Later among the works it cites.
“Consistent accelerated inference via confident adaptive transformers,”
Tal Schuster et al., · 2021
Later among the works it cites.
“LeeBERT: Learned early exit for BERT with cross-level optimization,”
Wei Zhu, · 2021
Later among the works it cites.
“GAML-BERT: Improving BERT early exiting by gradient aligned mutual learning,”
Wei Zhu, Xiaoling Wang, Yuan Ni, and Guotong Xie, · 2021
Later among the works it cites.
“Continually detection, rapidly react: Unseen rumors detection based on continual prompt-tuning,”
Yuhui Zuo, Wei Zhu, and Guoyong Cai, · 2022
Later among the works it cites.
“A survey on dynamic neural networks for natural language processing,”
Canwen et al. Xu, · 2022
Later among the works it cites.
“Pcee-bert: Accelerating bert inference via patient and confident early exiting,”
Zhen Zhang, Wei Zhu, Jinfan Zhang, Peng Wang, Rize Jin, and Tae-Sun Chung, · 2022
Later among the works it cites.
“Acf: Aligned contrastive finetuning for language and vision tasks,”
Wei Zhu, Peng Wang, Xiaoling Wang, Yuan Ni, and Guo Tong Xie, · 2023
Closest in time.