Optimal brain damage. In Advances in neural information processing systems . 598–605
Yann LeCun, John S Denker, and Sara A Solla. 1990 · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G Stork. 1993 · 1993
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation. In Advances in neural information processing systems . 1269–1277
Emily L Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Deep learning with limited numerical precision. In International Conference on Machine Learning . 1737–1746
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Original
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets. In International Conference on Learning Representations
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Convolutional neural networks with low-rank regularization
Original
Cheng Tai, Tong Xiao, Yi Zhang, Xiaogang Wang, et al · 2015
Earlier work this paper cites.
Face model compression by distilling knowledge from neurons. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 30
Ping Luo, Zhenyao Zhu, Ziwei Liu, Xiaogang Wang, and Xiaoou Tang. 2016 · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks. In European conference on computer vision . Springer, 525–542
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Deep learning with low precision by half-wave gaussian quantization. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5918–5926
Zhaowei Cai, Xiaodong He, Jian Sun, and Nuno Vasconcelos. 2017 · 2017
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming. In Proceedings of the IEEE International Conference on Computer Vision . 2736–2744
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
On compressing deep models by low rank and sparse decomposition. In IEEE Conference on Computer Vision and Pattern Recognition . 7370–7379
Xiyu Yu, Tongliang Liu, Xinchao Wang, and Dacheng Tao. 2017 · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Original
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2017 · 2017
Earlier work this paper cites.
Holistic cnn compression via low-rank decomposition with knowledge transfer
Shaohui Lin, Rongrong Ji, Chao Chen, Dacheng Tao, and Jiebo Luo. 2018 · 2018
Earlier work this paper cites.