Fetching the paper…
Reading the bibliography…
Knowledge distillation is a technique for improving the performance of a simple "student" model by replacing its one-hot training labels with a distribution over labels obtained from a complex "teacher" model.
Learning without forgetting
Z. Li and D. Hoiem · 1939
Earlier work this paper cites.
Probability inequalities for the sum of independent random variables
George Bennett · 1962
Earlier work this paper cites.
Elicitation of personal probabilities and expectations
Leonard J. Savage · 1971
Earlier work this paper cites.
A general method for comparing probability assessors
Mark J. Schervish · 1989
Earlier work this paper cites.
Extracting tree-structured representations of trained networks
Mark W. Craven and Jude W. Shavlik · 1995
Earlier work this paper cites.
Born again trees
Leo Breiman and Nong Shang · 1996
Earlier work this paper cites.
Learning to order things
William W. Cohen, Robert E. Schapire, and Yoram Singer · 1999
Earlier work this paper cites.
On the algorithmic implementation of multiclass kernel-based vector machines
Koby Crammer and Yoram Singer · 2002
Earlier work this paper cites.
Understanding and improving knowledge distillation
Jiaxi Tang, Rakesh Shivanna, Zhe Zhao, Dong Lin, Anima Singh, Ed H. Chi, and Sagar Jain · 2002
Earlier work this paper cites.
Log-linear models for label ranking
Ofer Dekel, Christopher D. Manning, and Yoram Singer · 2003
Earlier work this paper cites.
Loss functions for binary class probability estimation and classification: Structure and applications
Andreas Buja, Werner Stuetzle, and Yi Shen · 2005
Earlier work this paper cites.
Model compression
Cristian Bucilǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Multilabel classification via calibrated label ranking
Johannes Fürnkranz, Eyke Hüllermeier, Eneldo Loza Mencía, and Klaus Brinker · 2008
Earlier work this paper cites.
Empirical bernstein bounds and sample-variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
The P-Norm Push: A simple convex ranking algorithm that concentrates at the top of the list
Cynthia Rudin · 2009
Earlier work this paper cites.
Ranking with ordered weighted pairwise classification
Nicolas Usunier, David Buffoni, and Patrick Gallinari · 2009
Earlier work this paper cites.
Label Ranking Algorithms: A Survey , pages 45–64
Shankar Vembu and Thomas Gärtner · 2011
Earlier work this paper cites.
Hidden factors and hidden topics: Understanding rating dimensions with review text
Julian McAuley and Jure Leskovec · 2013
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
Jian Xue, Jinyu Li, and Yifan Gong · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Ranking via robust binary classification
Hyokun Yun, Parameswaran Raman, and S. V. N. Vishwanathan · 2014
Cited alongside, same era.
Sparse local embeddings for extreme multi-label classification
Kush Bhatia, Himanshu Jain, Purushottam Kar, Manik Varma, and Prateek Jain · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Cited alongside, same era.
Large-margin softmax loss for convolutional neural networks
Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang · 2016
Cited alongside, same era.
Unifying distillation and privileged information
D. Lopez-Paz, B. Schölkopf, L. Bottou, and V. Vapnik · 2016
Cited alongside, same era.
Distillation as a defense to adversarial perturbations against deep neural networks
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Later among the works it cites.
An alternative cross entropy loss for learning-to-rank
Sebastian Bruch · 2019
Later among the works it cites.
An analysis of the softmax cross entropy loss for learning-to-rank with binary relevance
Sebastian Bruch, Xuanhui Wang, Michael Bendersky, and Marc Najork · 2019
Later among the works it cites.
Learning imbalanced datasets with label-distribution-aware margin loss
Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Aréchiga, and Tengyu Ma · 2019
Later among the works it cites.
Distillation ≈ \approx early stopping? harvesting dark knowledge utilizing anisotropic information retrieval for overparameterized neural network, 2019
Bin Dong, Jikai Hou, Yiping Lu, and Zhihua Zhang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami · 2016
Cited alongside, same era.
Policy distillation
Andrei A. Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Recurrent neural network training with dark knowledge transfer
Zhiyuan Tang, Dong Wang, and Zhiyong Zhang · 2016
Cited alongside, same era.
Patient-driven privacy control through generalized distillation
Z. B. Celik, D. Lopez-Paz, and P. McDaniel · 2017
Cited alongside, same era.
Sobolev training for neural networks
Wojciech M. Czarnecki, Simon Osindero, Max Jaderberg, Grzegorz Swirszcz, and Razvan Pascanu · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Hypothesis set stability and generalization
Dylan J Foster, Spencer Greenberg, Satyen Kale, Haipeng Luo, Mehryar Mohri, and Karthik Sridharan · 2019
Later among the works it cites.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2019
Later among the works it cites.
Slice: Scalable linear extreme classifiers trained on 100 million labels for related searches
Himanshu Jain, Venkatesh Balasubramanian, Bhanu Chunduri, and Manik Varma · 2019
Later among the works it cites.
Striking the right balance with uncertainty
S. Khan, M. Hayat, S. W. Zamir, J. Shen, and L. Shao · 2019
Later among the works it cites.
Overfitting of neural nets under class imbalance: Analysis and improvements for segmentation
Zeju Li, Konstantinos Kamnitsas, and Ben Glocker · 2019
Later among the works it cites.
Search to distill: Pearls are everywhere but not the eyes, 2019
Yu Liu, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli, Yukun Zhu, Bradley Green, and Xiaogang Wang · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton · 2019
Later among the works it cites.
Zero-shot knowledge distillation in deep networks
Gaurav Kumar Nayak, Konda Reddy Mopuri, Vaisakh Shaj, Venkatesh Babu Radhakrishnan, and Anirban Chakraborty · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Later among the works it cites.
Stochastic negative mining for learning with large output spaces
Sashank J. Reddi, Satyen Kale, Felix Yu, Daniel Holtmann-Rice, Jiecao Chen, and Sanjiv Kumar · 2019
Later among the works it cites.
Noise regularization for conditional density estimation, 2019
Jonas Rothfuss, Fabio Ferreira, Simon Boehm, Simon Walther, Maxim Ulrich, Tamim Asfour, and Andreas Krause · 2019
Later among the works it cites.
Self-training with noisy student improves imagenet classification, 2019
Qizhe Xie, Eduard Hovy, Minh-Thang Luong, and Quoc V. Le · 2019
Later among the works it cites.
Training deep neural networks in generations: A more tolerant teacher educates better students
Chenglin Yang, Lingxi Xie, Siyuan Qiao, and Alan L. Yuille · 2019
Later among the works it cites.
Self-distillation amplifies regularization in hilbert space, 2020
Hossein Mobahi, Mehrdad Farajtabar, and Peter L. Bartlett · 2020
Closest in time.