Fetching the paper…
Reading the bibliography…
Knowledge distillation has been widely used to compress existing deep learning models while preserving the performance on a wide range of applications.
“Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,”
John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, and David S Pallett, · 1993
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
Combining pattern classifiers: methods and algorithms
Ludmila I Kuncheva, · 2014
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, · 2015
Earlier work this paper cites.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Sequence-level knowledge distillation,”
Yoon Kim and Alexander M Rush, · 2016
Earlier work this paper cites.
“Distilling knowledge from ensembles of neural networks for speech recognition.,”
Yevgen Chebotar and Austin Waters, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Earlier work this paper cites.
“Learning efficient object detection models with knowledge distillation,”
Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker, · 2017
Earlier work this paper cites.
“Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy,”
Asit Mishra and Debbie Marr, · 2017
Cited alongside, same era.
“Knowledge distillation across ensembles of multilingual models for low-resource languages,”
Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nussbaum-Thom, and Andrew Rosenberg, · 2017
Cited alongside, same era.
“Ensemble distillation for neural machine translation,”
Markus Freitag, Yaser Al-Onaizan, and Baskaran Sankaran, · 2017
Cited alongside, same era.
“Efficient knowledge distillation from an ensemble of teachers.,”
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas, Jia Cui, and Bhuvana Ramabhadran, · 2017
Cited alongside, same era.
“Joint ctc-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
“Tinybert: Distilling bert for natural language understanding,”
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu, · 2019
Later among the works it cites.
“Patient knowledge distillation for bert model compression,”
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu, · 2019
Later among the works it cites.
“Guiding ctc posterior spike timings for improved posterior fusion and knowledge distillation,”
Gakuto Kurata and Kartik Audhkhasi, · 2019
Later among the works it cites.
“Knowledge distillation using output errors for self-attention end-to-end models,”
Ho-Gyeong Kim, Hwidong Na, Hoshik Lee, Jihyun Lee, Tae Gyoon Kang, Min-Joong Lee, and Young Sang Choi, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Model compression via distillation and quantization,”
Antonio Polino, Razvan Pascanu, and Dan Alistarh, · 2018
Cited alongside, same era.
“Multi-label image classification via knowledge distillation from weakly-supervised detection,”
Yongcheng Liu, Lu Sheng, Jing Shao, Junjie Yan, Shiming Xiang, and Chunhong Pan, · 2018
Cited alongside, same era.
“An investigation of a knowledge distillation method for ctc acoustic models,”
Ryoichi Takashima, Sheng Li, and Hisashi Kawai, · 2018
Cited alongside, same era.
“Knowledge distillation for sequence model.,”
Mingkun Huang, Yongbin You, Zhehuai Chen, Yanmin Qian, and Kai Yu, · 2018
Cited alongside, same era.
“Private model compression via knowledge distillation,”
Ji Wang, Weidong Bao, Lichao Sun, Xiaomin Zhu, Bokai Cao, and S Yu Philip, · 2019
Cited alongside, same era.
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber, · 2019
Later among the works it cites.
“The pytorch-kaldi speech recognition toolkit,”
Mirco Ravanelli, Titouan Parcollet, and Yoshua Bengio, · 2019
Later among the works it cites.
“RWTH ASR Systems for LibriSpeech: Hybrid vs Attention,”
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“A federated approach in training acoustic models,”
Dimitrios Dimitriadis, Kenichi Kumatani, Robert Gmyr, Yashesh Gaur, and Sefik Emre Eskimez, · 2020
Closest in time.
“Federated learning: Challenges, methods, and future directions,”
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith, · 2020
Closest in time.
“Speechbrain: A general-purpose speech toolkit,”
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, et al., · 2021
Closest in time.