Fetching the paper…
Reading the bibliography…
Most teacher-student frameworks based on knowledge distillation (KD) depend on a strong congruent constraint on instance level.
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu · 1904
Earlier work this paper cites.
Model compression
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
E. L. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
M. Jaderberg, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Discovering structure in high-dimensional data through correlation explanation
G. Ver Steeg and A. Galstyan · 2014
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
S. Han, J. Pool, J. Tran, and W. Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Cited alongside, same era.
Bilinear cnn models for fine-grained visual recognition
T.-Y. Lin, A. RoyChowdhury, and S. Maji · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Ms-celeb-1m: A dataset and benchmark for large-scale face recognition
Y. Guo, L. Zhang, Y. Hu, X. He, and J. Gao · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
I. Hubara, M. Courbariaux, D. Soudry, E. Y. Ran, and Y. Bengio · 2016
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Later among the works it cites.
Mimicking very efficient network for object detection
Q. Li, S. Jin, and J. Yan · 2017
Later among the works it cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Later among the works it cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
X. Zhang, X. Zhou, M. Lin, and J. Sun · 2017
Later among the works it cites.
Large scale distributed neural network training through online distillation
R. Anil, G. Pereyra, A. Passos, R. Ormandi, G. E. Dahl, and G. E. Hinton · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The megaface benchmark: 1 million faces for recognition at scale
I. Kemelmachershlizerman, S. M. Seitz, D. Miller, and E. Brossard · 2016
Cited alongside, same era.
Face model compression by distilling knowledge from neurons
P. Luo, Z. Zhu, Z. Liu, X. Wang, X. Tang, et al · 2016
Cited alongside, same era.
Pruning convolutional neural networks for resource efficient inference
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz · 2016
Cited alongside, same era.
Deep model compression: Distilling knowledge from noisy teachers
B. B. Sau and V. N. Balasubramanian · 2016
Cited alongside, same era.
Do deep convolutional nets really need to be deep and convolutional?
G. Urban, K. J. Geras, S. E. Kahou, O. Aslan, S. Wang, R. Caruana, A. Mohamed, M. Philipose, and M. Richardson · 2016
Cited alongside, same era.
Quantized convolutional neural networks for mobile devices
J. Wu, L. Cong, Y. Wang, Q. Hu, and J. Cheng · 2016
Cited alongside, same era.
Arcface: Additive angular margin loss for deep face recognition
J. Deng, J. Guo, and S. Zafeiriou · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2018
Later among the works it cites.
Improving knowledge distillation with supporting adversarial samples
B. Heo, M. Lee, S. Yun, and J. Y. Choi · 2018
Later among the works it cites.
Knowledge distillation with adversarial samples supporting decisionboundary
B. Heo, M. Lee, S. Yun, and J. Y. Choi · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Later among the works it cites.
The devil of face recognition is in the noise
F. Wang, L. Chen, C. Li, S. Huang, Y. Chen, C. Qian, and C. C. Loy · 2018
Later among the works it cites.
Person transfer gan to bridge domain gap for person re-identification
L. Wei, S. Zhang, W. Gao, and Q. Tian · 2018
Later among the works it cites.
Training shallow and thin networks for acceleration via knowledge distillation with conditional adversarial networks
Z. Xu, Y.-C. Hsu, and J. Huang · 2018
Later among the works it cites.