Fetching the paper…
Reading the bibliography…
The pre-training models such as BERT have achieved great results in various natural language processing problems.
Distilling task-specific knowledge from bert into simple neural networks
Tang, R.; Lu, Y.; Liu, L.; Mou, L.; Vechtomova, O.; and Lin, J. 2019 · 1903
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Turc, I.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 1908
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2019 · 1909
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Deep learning face representation by joint identification-verification
Sun, Y.; Chen, Y.; Wang, X.; and Tang, X. 2014 · 1996
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020 · 2002
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W.; Wei, F.; Dong, L.; Bao, H.; Yang, N.; and Zhou, M. 2020 · 2002
Earlier work this paper cites.
Pre-trained models for natural language processing: A survey
Qiu, X.; Sun, T.; Xu, Y.; Shao, Y.; Dai, N.; and Huang, X. 2020 · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Hadsell, R.; Chopra, S.; and LeCun, Y. 2006 · 2006
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge
Bentivogli, L.; Clark, P.; Dagan, I.; and Giampiccolo, D. 2009 · 2009
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
Weinberger, K. Q.; and Saul, L. K. 2009 · 2009
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T.; Chen, K.; Corrado, G.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Gong, Y.; Liu, L.; Yang, M.; and Bourdev, L. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana, R.; Nicolas, B.; Ebrahimi, K. S.; Antoine, C.; Carlo, G.; and Yoshua, B. 2015 · 2015
Cited alongside, same era.
Han, S.; Mao, H.; and Dally, W. J. 2015 · 2015
Cited alongside, same era.
Learning both weights and connections for efficient neural network
Han, S.; Pool, J.; Tran, J.; and Dally, W. 2015 · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Cited alongside, same era.
Facenet: A unified embedding for face recognition and clustering
Schroff, F.; Kalenichenko, D.; and Philbin, J. 2015 · 2015
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Later among the works it cites.
Multilingual Neural Machine Translation with Knowledge Distillation
Tan, X.; Ren, Y.; He, D.; Qin, T.; Zhao, Z.; and Liu, T.-Y. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
MASS: Masked Sequence to Sequence Pre-training for Language Generation
Song, K.; Tan, X.; Qin, T.; Lu, J.; and Liu, T.-Y. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, T.; Goodfellow, I. J.; and Shlens, J. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Cited alongside, same era.
Improved deep metric learning with multi-class n-pair loss objective
Sohn, K. 2016 · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L. 2017 · 2017
Cited alongside, same era.
Adversarial Training Methods for Semi-Supervised Text Classification
Miyato, T.; Dai, A. M.; and Goodfellow, I. J. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A.; Nangia, N.; and Bowman, S. R. 2017 · 2017
Cited alongside, same era.
Patient Knowledge Distillation for BERT Model Compression
Sun, S.; Cheng, Y.; Gan, Z.; and Liu, J. 2019 · 2019
Later among the works it cites.
Contrastive Representation Distillation
Tian, Y.; Krishnan, D.; and Isola, P. 2019 · 2019
Later among the works it cites.
Small and Practical BERT Models for Sequence Labeling
Tsai, H.; Riesa, J.; Johnson, M.; Arivazhagan, N.; Li, X.; and Archer, A. 2019 · 2019
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019a · 2019
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A.; Singh, A.; and Bowman, S. R. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R. R.; and Le, Q. V. 2019 · 2019
Later among the works it cites.
Knowledge Distillation from Internal Representations
Aguilar, G.; Ling, Y.; Zhang, Y.; Yao, B.; Fan, X.; and Guo, C. 2020 · 2020
Closest in time.
Momentum contrast for unsupervised visual representation learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Closest in time.
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
Sun, Z.; Yu, H.; Song, X.; Liu, R.; Yang, Y.; and Zhou, D. 2020 · 2020
Closest in time.
Contrastive Representation Distillation
Tian, Y.; Krishnan, D.; and Isola, P. 2020 · 2020
Closest in time.
Heterogeneous graph attention network
Wang, X.; Ji, H.; Shi, C.; Wang, B.; Ye, Y.; Cui, P.; and Yu, P. S. 2019b · 2032
Closest in time.