Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) is the process of transferring knowledge from a large model to a small one.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
A general class of coefficients of divergence of one distribution from another
S. M. Ali and S. D. Silvey. 1966 · 1966
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla. 1989 · 1989
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Christopher M. Bishop. 2006 · 2006
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Pre-trained summarization distillation
Sam Shleifer and Alexander M Rush. 2020 · 2010
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
Findings of the 2016 Conference on Machine Translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
Rémi Lebret, David Grangier, and Michael Auli. 2016 · 2016
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Earlier work this paper cites.
f f -divergence inequalities
Igal Sason and Sergio Verdú. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
SeqGAN: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2017 · 2017
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Cited alongside, same era.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Efficient contextualized representation: Language model pruning for sequence labeling
Liyuan Liu, Xiang Ren, Jingbo Shang, Xiaotao Gu, Jian Peng, and Jiawei Han. 2018 · 2018
Cited alongside, same era.
Learning sparse neural networks through L 0 {L}_{0} regularization
Christos Louizos, Max Welling, and Diederik P Kingma. 2018 · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
ENGINE: Energy-based inference networks for non-autoregressive machine translation
Lifu Tu, Richard Yuanzhe Pang, Sam Wiseman, and Kevin Gimpel. 2020 · 2020
Later among the works it cites.
Model compression with two-stage multi-teacher knowledge distillation for web question answering system
Ze Yang, Linjun Shou, Ming Gong, Wutao Lin, and Daxin Jiang. 2020 · 2020
Later among the works it cites.
Dreaming to distill: Data-free knowledge transfer via DeepInversion
Hongxu Yin, Pavlo Molchanov, Jose M. Alvarez, Zhizhong Li, Arun Mallya, Derek Hoiem, Niraj K. Jha, and Jan Kautz. 2020 · 2020
Later among the works it cites.
Bridging maximum likelihood and adversarial learning via α \alpha -divergence
Miaoyun Zhao, Yulai Cong, Shuyang Dai, and Lawrence Carin. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Cited alongside, same era.
Findings of the 2019 Conference on Machine Translation (WMT19)
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019 · 2019
Cited alongside, same era.
Patient knowledge distillation for BERT model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Cited alongside, same era.
Why do neural dialog systems generate short and meaningless replies? A comparison between dialog and translation
Bolin Wei, Shuai Lu, Lili Mou, Hao Zhou, Pascal Poupart, Ge Li, and Zhi Jin. 2019 · 2019
Cited alongside, same era.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Cited alongside, same era.
Layer-wise model pruning based on mutual information
Chun Fan, Jiwei Li, Tianwei Zhang, Xiang Ao, Fei Wu, Yuxian Meng, and Xiaofei Sun. 2021 · 2021
Later among the works it cites.
Mosaicking to distill: Knowledge distillation from out-of-domain data
Gongfan Fang, Yifan Bao, Jie Song, Xinchao Wang, Donglin Xie, Chengchao Shen, and Mingli Song. 2021 · 2021
Later among the works it cites.
Improving task-agnostic BERT distillation with layer mapping search
Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2021 · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Later among the works it cites.
DART: Open-domain structured data record to text generation
Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani. 2021 · 2021
Later among the works it cites.
One teacher is enough? Pre-trained language model distillation from multiple teachers
Chuhan Wu, Fangzhao Wu, and Yongfeng Huang. 2021 · 2021
Later among the works it cites.
Learning noise transition matrix from only noisy labels via total variation regularization
Yivan Zhang, Gang Niu, and Masashi Sugiyama. 2021 · 2021
Later among the works it cites.
Commonsense-focused dialogues for response generation: An empirical study
Pei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, and Dilek Hakkani-Tur. 2021 · 2021
Later among the works it cites.
Non-autoregressive translation with layer-wise prediction and deep supervision
Chenyang Huang, Hao Zhou, Osmar R. Zaïane, Lili Mou, and Lei Li. 2022 · 2022
Later among the works it cites.
What makes data-to-text generation hard for pretrained language models?
Moniba Keymanesh, Adrian Benton, and Mark Dredze. 2022 · 2022
Later among the works it cites.
An unsupervised multiple-task and multiple-teacher model for cross-lingual named entity recognition
Zhuoran Li, Chunming Hu, Xiaohui Guo, Junfan Chen, Wenyi Qin, and Richong Zhang. 2022 · 2022
Later among the works it cites.
One reference is not enough: Diverse distillation with reference selection for non-autoregressive translation
Chenze Shao, Xuanfu Wu, and Yang Feng. 2022 · 2022
Later among the works it cites.
Sparse MLP for image recognition: Is self-attention really necessary?
Chuanxin Tang, Yucheng Zhao, Guangting Wang, Chong Luo, Wenxuan Xie, and Wenjun Zeng. 2022 · 2022
Later among the works it cites.
An equal-size hard EM algorithm for diverse dialogue generation
Yuqiao Wen, Yongchang Hao, Yanshuai Cao, and Lili Mou. 2023 · 2023
Closest in time.