Fetching the paper…
Reading the bibliography…
Knowledge transfer is shown to be a very successful technique for training neural classifiers: together with the ground truth data, it uses the "privileged information" (PI) obtained by a "teacher" network to train a "student" network.
On the Theory of Learning with Privileged Information. In Proceedings of the 23rd International Conference on Neural Information Processing Systems - Volume 2 (Vancouver, British Columbia, Canada) (NIPS’10) . Curran Associates Inc., Red Hook, NY, USA, 1894–1902
Dmitry Pechyony and Vladimir Vapnik. 2010 · 1902
Earlier work this paper cites.
Visual relationship detection with internal and external linguistic knowledge distillation. In Proceedings of the IEEE international conference on computer vision . 1974–1982
Ruichi Yu, Ang Li, Vlad I Morariu, and Larry S Davis. 2017 · 1982
Earlier work this paper cites.
Using the Nyström Method to Speed Up Kernel Machines
Christopher K. I. Williams and Matthias Seeger. 2001 · 2001
Earlier work this paper cites.
A new learning paradigm: Learning using privileged information
Vladimir Vapnik and Akshay Vashist. 2009 · 2009
Earlier work this paper cites.
Multiple Kernel Learning Algorithms
Mehmet Gönen and Ethem Alpayd. 2011 · 2011
Earlier work this paper cites.
Algorithms for Learning Kernels Based on Centered Alignment
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh. 2012 · 2012
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Sequence-Level Knowledge Distillation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Austin, Texas, 1317–1327
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Unifying distillation and privileged information. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik. 2016 · 2016
Cited alongside, same era.
Learning efficient object detection models with knowledge distillation. In Advances in Neural Information Processing Systems . 742–751
Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker. 2017 · 2017
Cited alongside, same era.
Knowledge transfer in SVM and neural networks
Vladimir Vapnik and Rauf Izmailov. 2017 · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 4133–4141
Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim. 2017 · 2017
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA . 322–332
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang. 2019 · 2019
Later among the works it cites.
Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8-14 December 2019, Vancouver, BC, Canada . 10835–10845
Yuan Cao and Quanquan Gu. 2019 · 2019
Later among the works it cites.
Width Provably Matters in Optimization for Deep Linear Neural Networks. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA . 1655–1664
Simon S. Du and Wei Hu. 2019 · 2019
Later among the works it cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh. 2018 · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clement Hongler. 2018 · 2018
Cited alongside, same era.
A novel Enhanced Collaborative Autoencoder with knowledge distillation for top-N recommender systems
Yiteng Pan, Fazhi He, and Haiping Yu. 2019 · 2018
Cited alongside, same era.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari. 2019 · 2019
Later among the works it cites.
Towards Understanding Knowledge Distillation. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, Long Beach, California, USA, 5142–5151
Mary Phuong and Christoph Lampert. 2019 · 2019
Later among the works it cites.
Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous Control. In 34th Conference on Neural Information Processing Systems (NeurIPS 2020)
Zhiyuan Xu, Kun Wu, Zhengping Che, Jian Tang, and Jieping Ye. 2020 · 2020
Closest in time.
Learning Using Privileged Information: Similarity Control and Knowledge Transfer
Vladimir Vapnik and Rauf Izmailov. 2015 · 2049
Closest in time.