Fetching the paper…
Reading the bibliography…
On-device Deep Neural Network (DNN) inference consumes significant computing resources and development efforts.
Least squares quantization in PCM
Stuart P. Lloyd. 1982 · 1982
Earlier work this paper cites.
Quantization
R.M. Gray and D.L. Neuhoff. 1998 · 1998
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA . IEEE Computer Society, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky. 2009 · 2009
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010 · 2010
Earlier work this paper cites.
Low rank matrix-valued Chernoff bounds and approximate matrix multiplication. In Proceedings of the twenty-second annual ACM-SIAM symposium on Discrete Algorithms . SIAM, 1422–1436
Avner Magen and Anastasios Zouzias. 2011 · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. 2011 · 2011
Earlier work this paper cites.
The German Traffic Sign Recognition Benchmark: A multi-class classification competition. In The 2011 International Joint Conference on Neural Networks . 1453–1460
Johannes Stallkamp, Marc Schlipsing, Jan Salmen, and Christian Igel. 2011 · 2011
Earlier work this paper cites.
Predicting parameters in deep learning
Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando De Freitas. 2013 · 2013
Earlier work this paper cites.
K-Means Hashing: An Affinity-Preserving Quantization Method for Learning Binary Compact Codes. In 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA, June 23-28, 2013 . IEEE Computer Society, 2938–2945
Kaiming He, Fang Wen, and Jian Sun. 2013 · 2013
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
Emily L Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. 2014 · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. 2015 · 2015
Earlier work this paper cites.
Quantization based Fast Inner Product Search
Ruiqi Guo, Sanjiv Kumar, Krzysztof Choromanski, and David Simcha. 2015 · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Mcdnn: An approximation-based execution framework for deep stream processing under resource constraints. In Proceedings of the 14th Annual International Conference on Mobile Systems, Applications, and Services . 123–136
Seungyeop Han, Haichen Shen, Matthai Philipose, Sharad Agarwal, Alec Wolman, and Arvind Krishnamurthy. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Categorical Reparameterization with Gumbel-Softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Earlier work this paper cites.
Deep quantization network for efficient image retrieval. In Proc. 13th AAAI Conf. Artif. Intell. 3457–3463
Cao Yue, M Long, J Wang, Zhu Han, and Q Wen. 2016 · 2016
Earlier work this paper cites.
DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou. 2016a · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou. 2016b · 2016
Cited alongside, same era.
Ternary neural networks for resource-efficient AI applications. In 2017 international joint conference on neural networks (IJCNN) . IEEE, 2547–2554
Hande Alemdar, Vincent Leroy, Adrien Prost-Boucle, and Frédéric Pétrot. 2017 · 2017
Cited alongside, same era.
Lcnn: Lookup-based convolutional neural network. In Proceedings of the IEEE conference on computer vision and pattern recognition . 7120–7129
Hessam Bagherinezhad, Mohammad Rastegari, and Ali Farhadi. 2017 · 2017
Cited alongside, same era.
Bolt: Accelerated data mining with fast vector compression. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 727–735
Davis W Blalock and John V Guttag. 2017 · 2017
Cited alongside, same era.
Once-for-all: Train one network and specialize it for efficient deployment
Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. 2019 · 2019
Later among the works it cites.
TensorFlow: An end-to-end open source machine learning platform
Google. 2019 · 2019
Later among the works it cites.
ONNX Runtime
Microsoft. 2019 · 2019
Later among the works it cites.
Spring Hill (NNP-I 1000) Intel’s Data Center Inference Chip. In 2019 IEEE Hot Chips 31 Symposium (HCS) . 1–12
Ofri Wechsler, Michael Behar, and Bharat Daga. 2019 · 2019
Later among the works it cites.
What is the state of neural network pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cachier: Edge-caching for recognition applications. In 2017 IEEE 37th international conference on distributed computing systems (ICDCS) . IEEE, 276–286
Utsav Drolia, Katherine Guo, Jiaqi Tan, Rajeev Gandhi, and Priya Narasimhan. 2017 · 2017
Cited alongside, same era.
Subic: A supervised, structured binary code for image search. In Proceedings of the IEEE international conference on computer vision . 833–842
Himalaya Jain, Joaquin Zepeda, Patrick Pérez, and Rémi Gribonval. 2017 · 2017
Cited alongside, same era.
In defense of product quantization
Benjamin Klein and Lior Wolf. 2017 · 2017
Cited alongside, same era.
Age progression/regression by conditional adversarial autoencoder. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5810–5818
Zhifei Zhang, Yang Song, and Hairong Qi. 2017 · 2017
Cited alongside, same era.
TVM: An automated end-to-end optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . 578–594
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Foggycache: Cross-device approximate computation reuse. In Proceedings of the 24th annual international conference on mobile computing and networking . 19–34
Peizhen Guo, Bo Hu, Rui Li, and Wenjun Hu. 2018 · 2018
Cited alongside, same era.
Potluck: Cross-application approximate deduplication for computation-intensive mobile applications. In Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems . 271–284
Peizhen Guo and Wenjun Hu. 2018 · 2018
Cited alongside, same era.
Yang Jiao, Liang Han, and Xin Long. 2020 · 2020
Later among the works it cites.
MACE. 2020
2020
Later among the works it cites.
Designing network design spaces. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10428–10436
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. 2020 · 2020
Later among the works it cites.
Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration. In 53rd Annual IEEE/ACM International Symposium on Microarchitecture, MICRO 2020, Athens, Greece, October 17-21, 2020 . IEEE, 88–101
Akshay Krishna Ramanathan, Gurpreet S. Kalsi, Srivatsa Srinivasa, Tarun Makesh Chandran, Kamlesh R. Pillai, Om Ji Omer, Vijaykrishnan Narayanan, and Sreenivas Subramoney. 2020 · 2020
Later among the works it cites.
And the Bit Goes Down: Revisiting the Quantization of Neural Networks. In ICLR 2020-Eighth International Conference on Learning Representations . 1–11
Pierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham, and Hervé Jégou. 2020 · 2020
Later among the works it cites.
LogicNets: Co-designed neural networks and circuits for extreme-throughput applications. In 2020 30th International Conference on Field-Programmable Logic and Applications (FPL) . IEEE, 291–297
Yaman Umuroglu, Yash Akhauri, Nicholas James Fraser, and Michaela Blott. 2020 · 2020
Later among the works it cites.
Multiplying Matrices Without Multiplying. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 992–1004
Davis Blalock and John Guttag. 2021 · 2021
Later among the works it cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. 2021 · 2021
Later among the works it cites.
Ten Lessons from Three Generations Shaped Google’s TPUv4i (ISCA ’21) . IEEE Press, 1–14
Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, and David Patterson. 2021 · 2021
Later among the works it cites.
Boosting mobile CNN inference through semantic memory. In Proceedings of the 29th ACM International Conference on Multimedia . 2362–2371
Yun Li, Chen Zhang, Shihao Han, Li Lyna Zhang, Baoqun Yin, Yunxin Liu, and Mengwei Xu. 2021 · 2021
Later among the works it cites.
Asymo: scalable and efficient deep-learning inference on asymmetric mobile cpus. In Proceedings of the 27th Annual International Conference on Mobile Computing and Networking . 215–228
Manni Wang, Shaohua Ding, Ting Cao, Yunxin Liu, and Fengyuan Xu. 2021 · 2021
Later among the works it cites.
ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization. In 55th IEEE/ACM International Symposium on Microarchitecture, MICRO 2022, Chicago, IL, USA, October 1-5, 2022 . IEEE, 1414–1433
Cong Guo, Chen Zhang, Jingwen Leng, Zihan Liu, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. 2022 · 2022
Later among the works it cites.
Romou: Rapidly generate high-performance tensor kernels for mobile gpus. In Proceedings of the 28th Annual International Conference on Mobile Computing And Networking . 487–500
Rendong Liang, Ting Cao, Jicheng Wen, Manni Wang, Yang Wang, Jianhua Zou, and Yunxin Liu. 2022 · 2022
Later among the works it cites.
Look-ups are not (yet) all you need for deep learning inference
Calvin McCarter and Nicholas Dronen. 2022 · 2022
Later among the works it cites.
Learnable lookup table for neural network quantization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 12423–12433
Longguang Wang, Xiaoyu Dong, Yingqian Wang, Li Liu, Wei An, and Yulan Guo. 2022 · 2022
Later among the works it cites.