Fetching the paper…
Reading the bibliography…
In the past few years, transformer-based pre-trained language models have achieved astounding success in both industry and academia.
Automatically constructing a corpus of sentential paraphrases
W. B. Dolan and C. Brockett · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. E. Dahl, and G. E. Hinton · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Y. Li, J. Yosinski, J. Clune, H. Lipson, and J. E. Hopcroft · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. S. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Z. Huang and N. Wang · 2017
Earlier work this paper cites.
Fixing weight decay regularization in adam
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J.-H. Bae, and J. Kim · 2017
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis · 2017
Earlier work this paper cites.
Senteval: An evaluation toolkit for universal sentence representations
A. Conneau and D. Kiela · 2018
Earlier work this paper cites.
Paraphrasing complex network: Network compression via factor transfer
J. Kim, S. Park, and N. Kwak · 2018
Earlier work this paper cites.
Dissecting contextual word embeddings: Architecture and representation
M. E. Peters, M. Neumann, L. Zettlemoyer, and W.-t. Yih · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
Variational information distillation for knowledge transfer
S. Ahn, S. X. Hu, A. C. Damianou, N. D. Lawrence, and Z. Dai · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Adaptive regularization of labels
Q. Ding, S. Wu, H. Sun, J. Guo, and S. Xia · 2019
Cited alongside, same era.
What does bert learn about the structure of language?
G. Jawahar, B. Sagot, and D. Seddah · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
S. Kornblith, M. Norouzi, H. Lee, and G. E. Hinton · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Meal v2: Boosting vanilla resnet-50 to 80%+ top-1 accuracy on imagenet without tricks
Z. Shen and M. Savvides · 2020
Later among the works it cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou · 2020
Later among the works it cites.
Contrastive representation distillation
Y. Tian, D. Krishnan, and P. Isola · 2020
Later among the works it cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
When does label smoothing help?
R. Müller, S. Kornblith, and G. E. Hinton · 2019
Cited alongside, same era.
Relational knowledge distillation
W. Park, D. Kim, Y. Lu, and M. Cho · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Cited alongside, same era.
Meal: Multi-model ensemble via adversarial learning
Z. Shen, Z. He, and X. Xue · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression
S. Sun, Y. Cheng, Z. Gan, and J. Liu · 2019
Cited alongside, same era.
Similarity-preserving knowledge distillation
F. Tung and G. Mori · 2019
Cited alongside, same era.
Well-read students learn better: The impact of student initialization on knowledge distillation
I. Turc, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Z. Wu, Z. Liu, J. Lin, Y. Lin, and S. Han · 2020
Later among the works it cites.
TextBrewer: An Open-Source Knowledge Distillation Toolkit for Natural Language Processing
Z. Yang, Y. Cui, Z. Chen, W. Che, T. Liu, S. Wang, and G. Hu · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization
L. Yuan, F. E. H. Tay, G. Li, T. Wang, and J. Feng · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Later among the works it cites.
Knowledge distillation: A survey
J. Gou, B. Yu, S. J. Maybank, and D. Tao · 2021
Later among the works it cites.
Npas: A compiler-aware framework of unified network pruning and architecture search for beyond real-time mobile acceleration
Z. Li, G. Yuan, W. Niu, P. Zhao, Y. Li, Y. Cai, X. Shen, Z. Zhan, Z. Kong, Q. Jin, Z. Chen, S. Liu, K. Yang, B. Ren, Y. Wang, and X. Lin · 2021
Later among the works it cites.
M6: A chinese multimodal pretrainer
J. Lin, R. Men, A. Yang, C. Zhou, M. Ding, Y. Zhang, P. Wang, A. Wang, L. Jiang, X. Jia, J. Zhang, J. Zhang, X. Zou, Z. Li, X. Q. Deng, J. Liu, J. Xue, H. Zhou, J. Ma, J. Yu, Y. Li, W. Lin, J. Zhou, J. ie Tang, and H. Yang · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Alp-kd: Attention-based layer projection for knowledge distillation
P. Passban, Y. Wu, M. Rezagholizadeh, and Q. Liu · 2021
Later among the works it cites.
How many layers and why? an analysis of the model depth in transformers
A. Simoulin and B. Crabbé · 2021
Later among the works it cites.
Minilmv2: Multi-head self-attention relation distillation for compressing pretrained transformers
W. Wang, H. Bao, S. Huang, L. Dong, and F. Wei · 2021
Later among the works it cites.
Revisiting few-sample bert fine-tuning
T. Zhang, F. Wu, A. Katiyar, K. Q. Weinberger, and Y. Artzi · 2021
Later among the works it cites.
Edgeformer: A parameter-efficient transformer for on-device seq2seq generation
T. Ge and F. Wei · 2022
Closest in time.
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang · 2022
Closest in time.
An iot system using deep learning to classify camera trap images on the edge
I. A. Zualkernan, S. Dhou, J. Judas, A. R. Sajun, B. R. Gomez, and L. A. Hussain · 2022
Closest in time.