Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models.
A theoretical analysis of contrastive unsupervised representation learning
Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak, Orestis Plevrakis, and Nikunj Saunshi. 2019 · 1902
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R. Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 1902
Earlier work this paper cites.
Distilling task-specific knowledge from bert into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 1905
Earlier work this paper cites.
Roberta: A robustly optimized fbertg pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 1907
Earlier work this paper cites.
Revealing the dark secrets of bert
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 1908
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 1908
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf. 2019a · 1910
Earlier work this paper cites.
Contrastive representation distillation
Y. Tian, D. Krishnan, and P. Isola. 2019 · 1910
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2002
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020b · 2002
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020b · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Cited alongside, same era.
Dinghan Shen, Mingzhi Zheng, Yelong Shen, Yanru Qu, and Weizhu Chen. 2020 · 2009
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Cited alongside, same era.
Why skip if you can combine: A simple knowledge distillation technique for intermediate layers
Yimeng Wu, Peyman Passban, Mehdi Rezagholizade, and Qun Liu. 2020 · 2010
Cited alongside, same era.
Towards zero-shot knowledge distillation for natural language processing
Ahmad Rashid, Vasileios Lioutas, Abbas Ghaddar, and Mehdi Rezagholizadeh. 2020 · 2012
What Does BERT Learn about the Structure of Language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. 2019 · 2019
Later among the works it cites.
Freelb: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
Xtremedistil: Multi-stage distillation for massive multilingual models
Subhabrata Mukherjee and Ahmed Hassan Awadallah. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Minilmv2: Multi-head self-attention relation distillation for compressing pretrained transformers
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2020a · 2012
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff. Dean. 2014 · 2014
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Cited alongside, same era.
Adversarial training methods for semi-supervised text classification
Takeru Miyato, Andrew M Dai, and Ian Goodfellow. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018 · 2018
Cited alongside, same era.
Hao Fu, Shaojun Zhou an Qihong Yang, Junjie Tang an Guiquan Liu, Kaikui Liu, and Xiaolong Li. 2021 · 2021
Later among the works it cites.
Generate, annotate, and learn: Generative models advance self-training and knowledge distillation
Xuanli He, Islam Nassar, Jamie Kiros, Gholamreza Haffari, and Mohammad Norouzi. 2021 · 2021
Later among the works it cites.
Not far away, not so close: Sample efficient nearest neighbour data augmentation via minimax
Ehsan Kamalloo, Mehdi Rezagholizadeh, Peyman Passban, and Ali Ghodsi. 2021 · 2021
Later among the works it cites.
Tianda Li, Ahmad Rashid, Aref Jafari, Pranav Sharma, Ali Ghodsi, and Mehdi Rezagholizadeh. 2021 · 2021
Later among the works it cites.
Alp-kd: Attention-based layer projection for knowledge distillation
Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, and Qun Liu. 2021 · 2021
Later among the works it cites.
Coda: Contrast-enhanced and diversity promoting data augmentation for natural language understanding
Yanru Qu, Dinghan Shen, Yelong Shen, Sandra Sajeev, Jiawei Han, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
MATE-KD: Masked adversarial TExt, a companion to knowledge distillation
Ahmad Rashid, Vasileios Lioutas, and Mehdi Rezagholizadeh. 2021 · 2021
Later among the works it cites.
Pro-kd: Progressive distillation by following the footsteps of the teacher
Mehdi Rezagholizadeh, Aref Jafari, Puneeth Salad, Pranav Sharma, Ali Saheb Pasand, and Ali Ghodsi. 2021 · 2021
Later among the works it cites.
Universal-kd: Attention-based output-grounded intermediate layer knowledge distillation
Yimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md Akmal Haidar, and Ali Ghodsi. 2021 · 2021
Later among the works it cites.
Ehsan Kamalloo, Mehdi Rezagholizadeh, and Ali Ghodsi. 2022 · 2022
Closest in time.