Fetching the paper…
Reading the bibliography…
Binarized Neural Network (BNN) removes bitwidth redundancy in classical CNN by using a single bit (-1/+1) for network parameters and intermediate representations, which has greatly reduced the off-chip data transfer and storage overhead.
NVIDIA CUDA Compute Unified Device Architecture Programming Guide
NVIDIA Corporation. 2007 · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton. 2009 · 2009
Earlier work this paper cites.
Min Lin, Qiang Chen, and Shuicheng Yan. 2013 · 2013
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
BinaryConnect: Training Deep Neural Networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. 2015 · 2015
Earlier work this paper cites.
Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks. In Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA ’15) . ACM, New York, NY, USA, 161–170
Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong. 2015 · 2015
Earlier work this paper cites.
YodaNN: An Ultra-Low Power Convolutional Neural Network Accelerator Based on Binary Weights. In 2016 IEEE Computer Society Annual Symposium on VLSI (ISVLSI) . 236–241
R. Andri, L. Cavigelli, D. Rossi, and L. Benini. 2016 · 2016
Earlier work this paper cites.
Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks. In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . 367–379
Y. Chen, J. Emer, and V. Sze. 2016 · 2016
Earlier work this paper cites.
BinaryNet: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
Matthieu Courbariaux and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
A high performance FPGA-based accelerator for large-scale convolutional neural networks. In 2016 26th International Conference on Field Programmable Logic and Applications (FPL) . 1–9
Huimin Li, Xitian Fan, Li Jiao, Wei Cao, Xuegong Zhou, and Lingli Wang. 2016 · 2016
Earlier work this paper cites.
Scalable and modularized RTL compilation of Convolutional Neural Networks onto FPGA. In 2016 26th International Conference on Field Programmable Logic and Applications (FPL) . 1–8
Yufei Ma, N. Suda, Yu Cao, J. Seo, and S. Vrudhula. 2016 · 2016
Earlier work this paper cites.
Going Deeper with Embedded FPGA Platform for Convolutional Neural Network. In Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA ’16) . ACM, New York, NY, USA, 26–35
Jiantao Qiu, Jie Wang, Song Yao, Kaiyuan Guo, Boxun Li, Erjin Zhou, Jincheng Yu, Tianqi Tang, Ningyi Xu, Sen Song, Yu Wang, and Huazhong Yang. 2016 · 2016
Cited alongside, same era.
XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016 · 2016
Cited alongside, same era.
14.6 A 1.42 TOPS/W deep convolutional neural network recognition processor for intelligent IoE systems. In 2016 IEEE International Solid-State Circuits Conference (ISSCC) . 264–265
J. Sim, J. Park, M. Kim, D. Bae, Y. Choi, and L. Kim. 2016 · 2016
Cited alongside, same era.
Throughput-Optimized OpenCL-based FPGA Accelerator for Large-Scale Convolutional Neural Networks. In Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA ’16) . ACM, New York, NY, USA, 16–25
Finn: A framework for fast, scalable binarized neural network inference. In Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 65–74
Yaman Umuroglu, Nicholas J Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, and Kees Vissers. 2017 · 2017
Later among the works it cites.
Accelerating binarized convolutional neural networks with software-programmable fpgas. In Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 15–24
Ritchie Zhao, Weinan Song, Wentao Zhang, Tianwei Xing, Jeng-Hau Lin, Mani Srivastava, Rajesh Gupta, and Zhiru Zhang. 2017 · 2017
Later among the works it cites.
Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, and Yurong Chen. 2017 · 2017
Later among the works it cites.
XNORBIN: A 95 TOp/s/W Hardware Accelerator for Binary Convolutional Neural Networks
Andrawes Al Bahou, Geethan Karunaratne, Renzo Andri, Lukas Cavigelli, and Luca Benini. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Naveen Suda, Vikas Chandra, Ganesh Dasika, Abinash Mohanty, Yufei Ma, Sarma Vrudhula, Jae-sun Seo, and Yu Cao. 2016 · 2016
Cited alongside, same era.
Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks. In 2016 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . 1–8
C. Zhang, Zhenman Fang, Peipei Zhou, Peichen Pan, and Jason Cong. 2016 · 2016
Cited alongside, same era.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou. 2016 · 2016
Cited alongside, same era.
Chenzhuo Zhu, Song Han, Huizi Mao, and William J. Dally. 2016 · 2016
Cited alongside, same era.
BRein Memory: A Single-Chip Binary/Ternary Reconfigurable in-Memory Deep Neural Network Accelerator Achieving 1.4 TOPS at 0.6 W
K. Ando, K. Ueyoshi, K. Orimo, H. Yonekawa, S. Sato, H. Nakahara, S. Takamaeda-Yamazaki, M. Ikebe, T. Asai, T. Kuroda, and M. Motomura. 2018 · 2017
Cited alongside, same era.
ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA. In Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA ’17) . ACM, New York, NY, USA, 75–84
Song Han, Junlong Kang, Huizi Mao, Yiming Hu, Xin Li, Yubin Li, Dongliang Xie, Hong Luo, Song Yao, Yu Wang, Huazhong Yang, and William (Bill) J. Dally. 2017 · 2017
Cited alongside, same era.
In-Datacenter Performance Analysis of a Tensor Processing Unit
Norman P. Jouppi, Cliff Young, Nishant Patil, and David A. Patterson et al. 2017 · 2017
Cited alongside, same era.
A kernel decomposition architecture for binary-weight convolutional neural networks. In Proceedings of the 54th Annual Design Automation Conference 2017 . ACM, 60
Hyeonuk Kim, Jaehyeong Sim, Yeongjae Choi, and Lee-Sup Kim. 2017 · 2017
Cited alongside, same era.
A 7.663-TOPS 8.2-W Energy-efficient FPGA Accelerator for Binary Convolutional Neural Networks. In FPGA . 290–291
Yixing Li, Zichuan Liu, Kai Xu, Hao Yu, and Fengbo Ren. 2017 · 2017
Cited alongside, same era.
Closest in time.
XNOR Neural Engine: a Hardware Accelerator IP for 21.6 fJ/op Binary Neural Network Inference
F. Conti, P. D. Schiavone, and L. Benini. 2018 · 2018
Closest in time.
UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition
Kartik Hegde, Jiyong Yu, Rohit Agrawal, Mengjia Yan, Michael Pellauer, and Christopher W Fletcher. 2018 · 2018
Closest in time.
FP-BNN: Binarized neural network on FPGA
Shuang Liang, Shouyi Yin, Leibo Liu, Wayne Luk, and Shaojun Wei. 2018b · 2018
Closest in time.
Computation Reuse in DNNs by Exploiting Input Similarity
Antonio González Marc Riera, Jose Maria Arnau. 2018 · 2018
Closest in time.
A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGA. In Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA ’18) . ACM, New York, NY, USA, 31–40
Hiroki Nakahara, Haruyoshi Yonekawa, Tomoya Fujii, and Shimpei Sato. 2018 · 2018
Closest in time.
Bit Fusion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Network. In 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) . 764–775
H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, V. Chandra, and H. Esmaeilzadeh. 2018 · 2018
Closest in time.
A Framework for Generating High Throughput CNN Implementations on FPGAs. In Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA ’18) . ACM, New York, NY, USA, 117–126
Hanqing Zeng, Ren Chen, Chi Zhang, and Viktor Prasanna. 2018 · 2018
Closest in time.
Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
Shilin Zhu, Xin Dong, and Hao Su. 2018 · 2018
Closest in time.