Fetching the paper…
Reading the bibliography…
Access to pre-trained models has recently emerged as a standard across numerous machine learning domains.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
Emily L. Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks, 2015
Alexander Mordvintsev, Christopher Olah, and Mike Tyka · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem · 2016
Earlier work this paper cites.
Distilling word embeddings: An encoding approach
Lili Mou, Ran Jia, Yan Xu, Ge Li, Lu Zhang, and Zhi Jin · 2016
Earlier work this paper cites.
Deep metric learning via lifted structured feature embedding
Hyun Oh Song, Yu Xiang, Stefanie Jegelka, and Silvio Savarese · 2016
Earlier work this paper cites.
Cnnpack: Packing convolutional neural networks in the frequency domain
Yunhe Wang, Chang Xu, Shan You, Dacheng Tao, and Chao Xu · 2016
Earlier work this paper cites.
Learning efficient object detection models with knowledge distillation
Guobin Chen, Wongun Choi, Xiang Yu, Tony X. Han, and Manmohan Chandraker · 2017
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang · 2017
Earlier work this paper cites.
Knowledge distillation for small-footprint highway networks
Liang Lu, Michelle Guo, and Steve Renals · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim · 2017
Cited alongside, same era.
Learning from multiple teacher networks
Shan You, Chang Xu, Chao Xu, and Dacheng Tao · 2017
Cited alongside, same era.
On compressing deep models by low rank and sparse decomposition
Xiyu Yu, Tongliang Liu, Xinchao Wang, and Dacheng Tao · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2019
Later among the works it cites.
Snapshot distillation: Teacher-student optimization in one generation
Chenglin Yang, Lingxi Xie, Chi Su, and Alan L. Yuille · 2019
Later among the works it cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma · 2019
Later among the works it cites.
Tinybert: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2020
Later among the works it cites.
Dreaming to distill: Data-free knowledge transfer via deepinversion
Hongxu Yin, Pavlo Molchanov, Jose M. Alvarez, Zhizhong Li, Arun Mallya, Derek Hoiem, Niraj K. Jha, and Jan Kautz · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Zagoruyko and Nikos Komodakis · 2017
Cited alongside, same era.
Super-convergence: Very fast training of neural networks using large learning rates
Leslie N. Smith and Nicholay Topin · 2018
Cited alongside, same era.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2018
Cited alongside, same era.
DAFL: Data-free learning of student networks
Hanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang, Chuanjian Liu, Boxin Shi, Chunjing Xu, Chao Xu, and Qi Tian · 2019
Cited alongside, same era.
Analyzing and improving representations with the soft nearest neighbor loss
Nicholas Frosst, Nicolas Papernot, and Geoffrey E. Hinton · 2019
Cited alongside, same era.
Learning lightweight lane detection cnns by self attention distillation
Yuenan Hou, Zheng Ma, Chunxiao Liu, and Chen Change Loy · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao · 2021
Later among the works it cites.
Robustness and diversity seeking data-free knowledge distillation
Pengchao Han, Jihong Park, Shiqiang Wang, and Yejun Liu · 2021
Later among the works it cites.
Vision transformer for small-size datasets
Seung Hoon Lee, Seunghyun Lee, and Byung Cheol Song · 2021
Later among the works it cites.
Data-free knowledge transfer: A survey
Yuang Liu, Wei Zhang, Jun Wang, and Jianyong Wang · 2021
Later among the works it cites.
A survey on curriculum learning
Xin Wang, Yudong Chen, and Wenwu Zhu · 2021
Later among the works it cites.
Cold diffusion: Inverting arbitrary image transforms without noise
Arpit Bansal, Eitan Borgnia, Hong-Min Chu, Jie S. Li, Hamid Kazemi, Furong Huang, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2022
Later among the works it cites.
Soft diffusion: Score matching for general corruptions
Giannis Daras, Mauricio Delbracio, Hossein Talebi, Alexandros G. Dimakis, and Peyman Milanfar · 2022
Later among the works it cites.
The forward-forward algorithm: Some preliminary investigations
Geoffrey Hinton · 2022
Later among the works it cites.
Discovering and overcoming limitations of noise-engineered data-free knowledge distillation
Piyush Raikwar and Deepak Mishra · 2022
Later among the works it cites.