Fetching the paper…
Reading the bibliography…
We characterize the statistical efficiency of knowledge transfer through $n$ samples from a teacher to a probabilistic student classifier with input space $\mathcal S$ over labels $\mathcal A$.
Contrastive representation distillation
Tian, Y · 1910
Earlier work this paper cites.
Jiang, H · 1911
Earlier work this paper cites.
Estimation des densités: risque minimax
Bretagnolle, J · 1979
Earlier work this paper cites.
A desicion-theoretic generalization of on-line learning and an application to boosting
Freund, Y · 1995
Earlier work this paper cites.
Born again trees
Breiman, L · 1996
Earlier work this paper cites.
Improving the accuracy and speed of support vector machines
Burges, C. J · 1996
Earlier work this paper cites.
Springer series in statistics
van der Vaart, A. W · 1996
Earlier work this paper cites.
How to use expert advice
Cesa-Bianchi, N · 1997
Earlier work this paper cites.
Assouad, fano, and le cam
Yu, B · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y · 2000
Earlier work this paper cites.
Asymptotic statistics
Van der Vaart, A. W · 2000
Earlier work this paper cites.
A short note on learning discrete distributions
Canonne, C. L · 2002
Earlier work this paper cites.
Müller, R · 2002
Earlier work this paper cites.
Understanding and improving knowledge distillation
Tang, J · 2002
Earlier work this paper cites.
Concentration inequalities for the missing mass and for histogram rule error
McAllester, D · 2003
Earlier work this paper cites.
Optimal rates of aggregation
Tsybakov, A. B · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P · 2004
Earlier work this paper cites.
Why distillation helps: a statistical perspective
Menon, A. K · 2005
Earlier work this paper cites.
Model compression
Buciluǎ, C · 2006
Earlier work this paper cites.
A survey on transfer learning in natural language processing
Alyafeai, Z · 2007
Earlier work this paper cites.
A coincidence-based test for uniformity given very sparsely sampled discrete data
Paninski, L · 2008
Earlier work this paper cites.
Toward the fundamental limits of imitation learning
Rajaraman, N · 2009
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Tsybakov, A. B · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S · 2010
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z · 2012
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, J · 2014
Earlier work this paper cites.
Learning small-size dnn with output-distribution-based criteria
Li, J · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
Lopez-Paz, D · 2015
Earlier work this paper cites.
A survey of transfer learning
Weiss, K · 2016
Earlier work this paper cites.
Zagoruyko, S · 2016
Cited alongside, same era.
Like what you like: Knowledge distill via neuron selectivity transfer
Huang, Z · 2017
Cited alongside, same era.
Data-free knowledge distillation for deep neural networks
Lopes, R. G · 2017
Cited alongside, same era.
Probability and computing: Randomization and probabilistic techniques in algorithms and data analysis
Mitzenmacher, M · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J · 2017
Cited alongside, same era.
Comparing kullback-leibler divergence and mean squared error loss in knowledge distillation
Kim, T · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P · 2021
Later among the works it cites.
Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks
Wang, L · 2021
Later among the works it cites.
Zero-shot knowledge distillation from a decision-based black-box model
Wang, Z · 2021
Later among the works it cites.
Near-optimal provable uniform convergence in offline policy evaluation for reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Born again neural networks
Furlanello, T · 2018
Cited alongside, same era.
Knowledge transfer with jacobian matching
Srinivas, S · 2018
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
Cho, J. H · 2019
Cited alongside, same era.
Probability: theory and examples
Durrett, R · 2019
Cited alongside, same era.
Zero-shot knowledge distillation in deep networks
Nayak, G. K · 2019
Cited alongside, same era.
Knockoff nets: Stealing functionality of black-box models
Orekondy, T · 2019
Cited alongside, same era.
Yin, M · 2021
Later among the works it cites.
Rethinking soft labels for knowledge distillation: A bias-variance tradeoff perspective
Zhou, H · 2021
Later among the works it cites.
Up to 100x faster data-free knowledge distillation
Fang, G · 2022
Later among the works it cites.
Black-box few-shot knowledge distillation
Nguyen, D · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Later among the works it cites.
Analysis of knowledge transfer in kernel regime
Panahi, A · 2022
Later among the works it cites.
Information theory: From coding to learning
Polyanskiy, Y · 2022
Later among the works it cites.
Better teacher better student: Dynamic prior knowledge for knowledge distillation
Qiu, Z · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A · 2022
Later among the works it cites.
Glm-130b: An open bilingual pre-trained model
Zeng, A · 2022
Later among the works it cites.
Decoupled knowledge distillation
Zhao, B · 2022
Later among the works it cites.
Gkd: Generalized knowledge distillation for auto-regressive sequence models
Agarwal, R · 2023
Closest in time.
The falcon series of language models: Towards open frontier models
Almazrouei, E · 2023
Closest in time.
Model card and evaluations for claude models
Anthropic · 2023
Closest in time.
Scaling laws for reward model overoptimization
Gao, L · 2023
Closest in time.
Boot: Data-free distillation of denoising diffusion models with bootstrapping
Gu, J · 2023
Closest in time.
On the impact of knowledge distillation for model interpretability
Han, H · 2023
Closest in time.
Supervision complexity and its role in knowledge distillation
Harutyunyan, H · 2023
Closest in time.
Lion: Adversarial distillation of closed-source large language model
Jiang, Y · 2023
Closest in time.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X · 2023
Closest in time.
Less is more: Task-aware layer-wise distillation for language model compression
Liang, C · 2023
Closest in time.
Liu, H · 2023
Closest in time.
Accessed: Sep. 28,2023
OpenAI · 2023
Closest in time.
Peng, B · 2023
Closest in time.
Knowledge distillation on graphs: A survey
Tian, Y · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L · 2023
Closest in time.
Learning using privileged information: similarity control and knowledge transfer
Vapnik, V · 2049
Closest in time.