Fetching the paper…
Reading the bibliography…
Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks.
Distilling task-specific knowledge from bert into simple neural networks
R. Tang, Y. Lu, L. Liu, L. Mou, O. Vechtomova, and J. Lin. 2019 · 1903
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov. 2019 · 1907
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
I. Turc, M. Chang, K. Lee, and K. Toutanova. 2019 · 1908
Earlier work this paper cites.
Q-bert: Hessian based ultra low precision quantization of bert
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer. 2019 · 1909
Earlier work this paper cites.
Pruning a bert-based question answering model
J. S. McCarley. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf. 2019 · 1910
Earlier work this paper cites.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat. 2019 · 1910
Earlier work this paper cites.
Compressing bert: Studying the effects of weight pruning on transfer learning
M. A. Gordon, K. Duh, and N. Andrews. 2020 · 2002
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou. 2020 · 2002
Earlier work this paper cites.
Bert-of-theseus: Compressing bert by progressive module replacing
C. Xu, W. Zhou, T. Ge, F. Wei, and M. Zhou. 2020 · 2002
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
L. Hou, L. Shang, X. Jiang, and Q. Liu. 2020 · 2004
Earlier work this paper cites.
Fastbert: a self-distilling bert with adaptive inference time
W. Liu, P. Zhou, Z. Zhao, Z. Wang, H. Deng, and Q. Ju. 2020 · 2004
Earlier work this paper cites.
Poor man’s bert: Smaller and faster transformer models
H. Sajjad, F. Dalvi, N. Durrani, and P. Nakov. 2020 · 2004
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou. 2020 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
W. B. Dolan and C. Brockett. 2005 · 2005
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
L. Bentivogli, P. Clark, I. Dagan, and D. Giampiccolo. 2009 · 2009
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts. 2013 · 2013
Cited alongside, same era.
Compressing deep convolutional networks using vector quantization
Y. Gong, L. Liu, M. Yang, and L. Bourdev. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning. 2014 · 2014
Cited alongside, same era.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio. 2014 · 2014
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning. 2019 · 2019
Closest in time.
Fine-tune bert with sparse self-attention mechanism
B. Cui, Y. Li, M. Chen, and Z. Zhang. 2019 · 2019
Closest in time.
Universal transformers
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser. 2019 · 2019
Closest in time.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova. 2019 · 2019
Closest in time.
Cognitive graph for multi-hop reading comprehension at scale
M. Ding, C. Zhou, Q. Chen, H. Yang, and J. Tang. 2019 · 2019
Closest in time.
Revealing the dark secrets of bert
O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S Han, J. Pool, J. Tran, and W. Dally. 2015 · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean. 2015 · 2015
Cited alongside, same era.
Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
S. Han, Mao H., and Dally W. J. 2016 · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Y. Kim and A. M. Rush. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. 2016 · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
D. Cer, M. Diab, E. Agirre, I. Lopez-Gazpio, and L. Specia. 2017 · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. 2017 · 2017
Cited alongside, same era.
A tensorized transformer for language modeling
X. Ma, P. Zhang, S. Zhang, N. Duan, Y. Hou, M. Zhou, and D. Song. 2019 · 2019
Closest in time.
Are sixteen heads really better than one?
P. Michel, O. Levy, and G. Neubig. 2019 · 2019
Closest in time.
Patient knowledge distillation for bert model compression
S. Sun, Y. Cheng, Z. Gan, and J. Liu. 2019 · 2019
Closest in time.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov. 2019 · 2019
Closest in time.
Neural network acceptability judgments
A. Warstadt, A. Singh, and S. R. Bowman. 2019 · 2019
Closest in time.
Conditional bert contextual augmentation
X. Wu, S. Lv, L. Zang, J. Han, and S. Hu. 2019 · 2019
Closest in time.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le. 2019 · 2019
Closest in time.
Electra: Pre-training text encoders as discriminators rather than generators
K. Clark, M. Luong, Q. V. Le, and C. D. Manning. 2020 · 2020
Closest in time.
Depth-adaptive transformer
M. Elbayad, J. Gu, E. Grave, and M. Auli. 2020 · 2020
Closest in time.
Reducing transformer depth on demand with structured dropout
Angela F., Edouard G., and Armand J. 2020 · 2020
Closest in time.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut. 2020 · 2020
Closest in time.