Fetching the paper…
Reading the bibliography…
Due to recent advances in digital technologies, and availability of credible data, an area of artificial intelligence, deep learning, has emerged, and has demonstrated its ability and effectiveness in solving complex learning problems not possible before.
J. MacQueen et al. , “Some methods for classification and analysis of multivariate observations,” in Proceedings of the fifth Berkeley symposium on mathematical statistics and probability , vol. 1, no. 14. Oakland, CA, USA, 1967, pp. 281–297
1967
Earlier work this paper cites.
S. Winograd, Arithmetic complexity of computations . Siam, 1980, vol. 33
1980
Earlier work this paper cites.
J. D. Dixon, “Asymptotically fast factorization of integers,” Mathematics of computation , vol. 36, no. 153, pp. 255–260, 1981
1981
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, no. 6088, p. 533, 1986
1986
Earlier work this paper cites.
E. A. Lee and D. G. Messerschmitt, “Synchronous data flow,” Proceedings of the IEEE , vol. 75, no. 9, pp. 1235–1245, 1987
1987
Earlier work this paper cites.
Rumelhart, David E and Hinton, Geoffrey E and Williams, Ronald J, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors , pp. 696–699, 1988
1988
Earlier work this paper cites.
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural computation , vol. 1, no. 4, pp. 541–551, 1989
1989
Earlier work this paper cites.
S. J. Hanson and L. Y. Pratt, “Comparing biases for minimal network construction with back-propagation,” in Advances in neural information processing systems , 1989, pp. 177–185
1989
Earlier work this paper cites.
J. E. Dayhoff, Neural network architectures: an introduction . Van Nostrand Reinhold New York, 1990
1990
Earlier work this paper cites.
Y. LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel, “Handwritten digit recognition with a back-propagation network,” in Advances in neural information processing systems , 1990, pp. 396–404
1990
Earlier work this paper cites.
Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal brain damage,” in Advances in neural information processing systems , 1990, pp. 598–605
1990
Earlier work this paper cites.
C. Van Loan, Computational frameworks for the fast Fourier transform . Siam, 1992, vol. 10
1992
Earlier work this paper cites.
B. Hassibi and D. G. Stork, “Second order derivatives for network pruning: Optimal brain surgeon,” in Advances in neural information processing systems , 1993, pp. 164–171
1993
Earlier work this paper cites.
D. F. Bacon, S. L. Graham, and O. J. Sharp, “Compiler transformations for high-performance computing,” ACM Computing Surveys (CSUR) , vol. 26, no. 4, pp. 345–420, 1994
1994
Earlier work this paper cites.
P. L. Montgomery, “A survey of modern integer factorization algorithms,” CWI quarterly , vol. 7, no. 4, pp. 337–365, 1994
1994
Earlier work this paper cites.
Y. LeCun and Y. Bengio, “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, p. 1995, 1995
1995
Earlier work this paper cites.
C. R. Reeves, Modern heuristic techniques for combinatorial problems. Advanced topics in computer science . Mc Graw-Hill, 1995
1995
Earlier work this paper cites.
C. F. Van Loan, “Matrix computations (johns hopkins studies in mathematical sciences),” 1996
1996
Earlier work this paper cites.
J. Cloutier, E. Cosatto, S. Pigeon, F. R. Boyer, and P. Y. Simard, “Vip: An FPGA-based processor for image processing and neural networks,” in Microelectronics for Neural Networks, 1996., Proceedings of Fifth International Conference on . IEEE, 1996, pp. 330–336
1996
Earlier work this paper cites.
A. Sato and K. Yamada, “Generalized learning vector quantization,” in Advances in neural information processing systems , 1996, pp. 423–429
1996
Earlier work this paper cites.
J. Villasenor and W. H. Mangione-Smith, “Configurable computing,” Scientific American , vol. 276, no. 6, pp. 66–71, 1997
1997
Earlier work this paper cites.
S. Lawrence, C. L. Giles, A. C. Tsoi, and A. D. Back, “Face recognition: A convolutional neural-network approach,” IEEE transactions on neural networks , vol. 8, no. 1, pp. 98–113, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
Y. LeCun, “The mnist database of handwritten digits,” http://yann. lecun. com/exdb/mnist/ , 1998
1998
Earlier work this paper cites.
R. C. Whaley and J. J. Dongarra, “Automatically tuned linear algebra software,” in Supercomputing, 1998. SC98. IEEE/ACM Conference on . IEEE, 1998, pp. 38–38
1998
Earlier work this paper cites.
J. C. Platt, “12 fast training of support vector machines using sequential minimal optimization,” Advances in kernel methods , pp. 185–208, 1999
1999
Earlier work this paper cites.
B. Bosi, G. Bois, and Y. Savaria, “Reconfigurable pipelined 2-d convolvers for fast digital signal processing,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 7, no. 3, pp. 299–308, 1999
1999
Earlier work this paper cites.
S. M. Sait and H. Youssef, Iterative computer algorithms with applications in engineering: solving combinatorial optimization problems . IEEE Computer Society Press, 1999
1999
Earlier work this paper cites.
P. J. Lisboa and E. C. Ifeachor, Artificial neural networks in biomedicine . Springer Science & Business Media, 2000
2000
Earlier work this paper cites.
D. F. Wolf, R. A. Romero, and E. Marques, “Using embedded processors in hardware models of artificial neural networks,” in V Simposio Brasileiro de automação inteligente, Brasil , 2001
2001
Earlier work this paper cites.
K. R. Nichols, M. A. Moussa, and S. M. Areibi, “Feasibility of floating-point arithmetic in FPGA based artificial neural networks,” in In CAINE . Citeseer, 2002
2002
Earlier work this paper cites.
K. Benkrid and S. Belkacemi, “Design and implementation of a 2d convolution core for video applications on FPGAs,” in Digital and Computational Video, 2002. DCV 2002. Proceedings. Third International Workshop on . IEEE, 2002, pp. 85–92
2002
Earlier work this paper cites.
P. Y. Simard, D. Steinkraus, and J. C. Platt, “Best practices for convolutional neural networks applied to visual document analysis,” in null . IEEE, 2003, p. 958
2003
Earlier work this paper cites.
K. Korekado, T. Morie, O. Nomura, T. Nakano, M. Matsugu, and A. Iwata, “An image filtering processor for face/object recognition using merged/mixed analog-digital architecture,” in VLSI Circuits, 2005. Digest of Technical Papers. 2005 Symposium on . IEEE, 2005, pp. 220–223
2005
Earlier work this paper cites.
P. D. McNelis, Neural networks in finance: gaining predictive edge in the market . Academic Press, 2005
2005
Earlier work this paper cites.
F. Cardells-Tormo and P.-L. Molinet, “Area-efficient 2-d shift-variant convolvers for FPGA-based digital image processing,” in Signal Processing Systems Design and Implementation, 2005. IEEE Workshop on . IEEE, 2005, pp. 209–213
2005
Earlier work this paper cites.
R. G. Gironés, R. C. Palero, J. C. Boluda, and A. S. Cortés, “FPGA implementation of a pipelined on-line backpropagation,” Journal of VLSI signal processing systems for signal, image and video technology , vol. 40, no. 2, pp. 189–213, 2005
2005
Earlier work this paper cites.
J. Mutch and D. G. Lowe, “Multiclass object recognition with sparse, localized features,” in Computer Vision and Pattern Recognition, 2006 IEEE Computer Society Conference on , vol. 1. IEEE, 2006, pp. 11–18
2006
Earlier work this paper cites.
U. Muller, J. Ben, E. Cosatto, B. Flepp, and Y. L. Cun, “Off-road obstacle avoidance through end-to-end learning,” in Advances in neural information processing systems , 2006, pp. 739–746
2006
Earlier work this paper cites.
A. R. Omondi and J. C. Rajapakse, FPGA implementations of neural networks . Springer, 2006, vol. 365
2006
Earlier work this paper cites.
K. Chellapilla, S. Puri, and P. Simard, “High performance convolutional neural networks for document processing,” in Tenth International Workshop on Frontiers in Handwriting Recognition . Suvisoft, 2006
2006
Earlier work this paper cites.
R. Hadsell, A. Erkan, P. Sermanet, J. Ben, K. Kavukcuoglu, U. Muller, and Y. LeCun, “A multi-range vision strategy for autonomous offroad navigation,” Proc. Robotics and Applications (RA’07) , vol. 1, no. 7, 2007
2007
Earlier work this paper cites.
O. Nomura and T. Morie, “Projection-field-type VLSI convolutional neural networks using merged/mixed analog-digital approach,” in International Conference on Neural Information Processing . Springer, 2007, pp. 1081–1090
2007
Earlier work this paper cites.
T. Serre, L. Wolf, S. Bileschi, M. Riesenhuber, and T. Poggio, “Robust object recognition with cortex-like mechanisms,” IEEE Transactions on Pattern Analysis & Machine Intelligence , no. 3, pp. 411–426, 2007
2007
Earlier work this paper cites.
H. Zhang, M. Xia, and G. Hu, “A multiwindow partial buffering scheme for FPGA-based 2-d convolvers,” IEEE Transactions on Circuits and Systems II: Express Briefs , vol. 54, no. 2, pp. 200–204, 2007
2007
Earlier work this paper cites.
A. W. Savich, M. Moussa, and S. Areibi, “The impact of arithmetic representation on implementing mlp-bp on FPGAs: A study,” IEEE transactions on neural networks , vol. 18, no. 1, pp. 240–252, 2007
2007
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: Deep neural networks with multitask learning,” in Proceedings of the 25th international conference on Machine learning . ACM, 2008, pp. 160–167
2008
Earlier work this paper cites.
P. W. Mirowski, Y. LeCun, D. Madhavan, and R. Kuzniecky, “Comparing SVM and convolutional networks for epileptic seizure prediction from intracranial EEG,” in Machine Learning for Signal Processing, 2008. MLSP 2008. IEEE Workshop on . IEEE, 2008, pp. 244–249
2008
Earlier work this paper cites.
M. C. Herbordt, Y. Gu, T. VanCourt, J. Model, B. Sukhwani, and M. Chiu, “Computing models for FPGA-based accelerators,” Computing in science & engineering , vol. 10, no. 6, pp. 35–45, 2008
2008
Earlier work this paper cites.
R. Collobert, C. Farabet, K. Kavukcuoglu et al. , “Torch,” in Workshop on Machine Learning Open Source Software, NIPS , vol. 76, 2008
2008
Earlier work this paper cites.
A. Beric, J. van Meerbergen, G. de Haan, and R. Sethuraman, “Memory-centric video processing,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 18, no. 4, pp. 439–452, 2008
2008
Earlier work this paper cites.
Y. Bengio et al. , “Learning deep architectures for ai,” Foundations and trends® in Machine Learning , vol. 2, no. 1, pp. 1–127, 2009
2009
Earlier work this paper cites.
P. Sermanet, R. Hadsell, M. Scoffier, M. Grimes, J. Ben, A. Erkan, C. Crudele, U. Miller, and Y. LeCun, “A multirange architecture for collision-free off-road robot navigation,” Journal of Field Robotics , vol. 26, no. 1, pp. 52–87, 2009
2009
Earlier work this paper cites.
R. Hadsell, P. Sermanet, J. Ben, A. Erkan, M. Scoffier, K. Kavukcuoglu, U. Muller, and Y. LeCun, “Learning long-range vision for autonomous off-road driving,” Journal of Field Robotics , vol. 26, no. 2, pp. 120–144, 2009
2009
Earlier work this paper cites.
A. Munshi, “The opencl specification,” in Hot Chips 21 Symposium (HCS), 2009 IEEE . IEEE, 2009, pp. 1–314
2009
Earlier work this paper cites.
C. Farabet, C. Poulet, J. Y. Han, and Y. LeCun, “Cnp: An FPGA-based processor for convolutional networks,” in Field Programmable Logic and Applications, 2009. FPL 2009. International Conference on . IEEE, 2009, pp. 32–37
2009
Earlier work this paper cites.
M. Sankaradas, V. Jakkula, S. Cadambi, S. Chakradhar, I. Durdanovic, E. Cosatto, and H. P. Graf, “A massively parallel coprocessor for convolutional neural networks,” in Application-specific Systems, Architectures and Processors, 2009. ASAP 2009. 20th IEEE International Conference on . IEEE, 2009, pp. 53–60
2009
Earlier work this paper cites.
H. P. Graf, S. Cadambi, V. Jakkula, M. Sankaradass, E. Cosatto, S. Chakradhar, and I. Dourdanovic, “A massively parallel digital learning processor,” in Advances in Neural Information Processing Systems , 2009, pp. 529–536
2009
Earlier work this paper cites.
S. Cadambi, I. Durdanovic, V. Jakkula, M. Sankaradass, E. Cosatto, S. Chakradhar, and H. P. Graf, “A massively parallel FPGA-based coprocessor for support vector machines,” in 2009 17th IEEE Symposium on Field Programmable Custom Computing Machines . IEEE, 2009, pp. 115–122
2009
Earlier work this paper cites.
F. Nasse, C. Thurau, and G. A. Fink, “Face detection using gpu-based convolutional neural networks,” in International Conference on Computer Analysis of Images and Patterns . Springer, 2009, pp. 83–90
2009
Earlier work this paper cites.
D. Grangier, L. Bottou, and R. Collobert, “Deep convolutional networks for scene parsing,” in ICML 2009 Deep Learning Workshop , vol. 3, no. 6. Citeseer, 2009, p. 109
2009
Earlier work this paper cites.
C. Farabet, C. Poulet, and Y. LeCun, “An FPGA-based stream processor for embedded real-time vision with convolutional networks,” in Computer Vision Workshops (ICCV Workshops), 2009 IEEE 12th International Conference on . IEEE, 2009, pp. 878–885
2009
Earlier work this paper cites.
S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Communications of the ACM , vol. 52, no. 4, pp. 65–76, 2009
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Citeseer, Tech. Rep., 2009
2009
Earlier work this paper cites.
J. Misra and I. Saha, “Artificial neural networks in hardware: A survey of two decades of progress,” Neurocomputing , vol. 74, no. 1-3, pp. 239–255, 2010
2010
Earlier work this paper cites.
R. Hameed, W. Qadeer, M. Wachs, O. Azizi, A. Solomatnikov, B. C. Lee, S. Richardson, C. Kozyrakis, and M. Horowitz, “Understanding sources of inefficiency in general-purpose chips,” in ACM SIGARCH Computer Architecture News , vol. 38, no. 3. ACM, 2010, pp. 37–47
2010
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proceedings of the 27th international conference on machine learning (ICML-10) , 2010, pp. 807–814
2010
Earlier work this paper cites.
J. E. Stone, D. Gohara, and G. Shi, “OpenCL: A parallel programming standard for heterogeneous computing systems,” Computing in science & engineering , vol. 12, no. 3, pp. 66–73, 2010
2010
Earlier work this paper cites.
S. Chakradhar, M. Sankaradas, V. Jakkula, and S. Cadambi, “A dynamically configurable coprocessor for convolutional neural networks,” in ACM SIGARCH Computer Architecture News , vol. 38, no. 3. ACM, 2010, pp. 247–257
2010
Earlier work this paper cites.
S. Cadambi, A. Majumdar, M. Becchi, S. Chakradhar, and H. P. Graf, “A programmable parallel accelerator for learning and classification,” in Proceedings of the 19th international conference on Parallel architectures and compilation techniques . ACM, 2010, pp. 273–284
2010
Earlier work this paper cites.
B. Bai, J. Weston, D. Grangier, R. Collobert, K. Sadamasa, Y. Qi, O. Chapelle, and K. Weinberger, “Learning to rank with (a lot of) word features,” Information retrieval , vol. 13, no. 3, pp. 291–314, 2010
2010
Cited alongside, same era.
C. Farabet, B. Martini, P. Akselrod, S. Talay, Y. LeCun, and E. Culurciello, “Hardware accelerated convolutional neural networks for synthetic vision systems,” in Circuits and Systems (ISCAS), Proceedings of 2010 IEEE International Symposium on . IEEE, 2010, pp. 257–260
2010
Cited alongside, same era.
J. Cong, B. Liu, S. Neuendorffer, J. Noguera, K. Vissers, and Z. Zhang, “High-level synthesis for fpgas: From prototyping to deployment,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 30, no. 4, pp. 473–491, 2011
2011
Cited alongside, same era.
A. Canis, J. Choi, M. Aldham, V. Zhang, A. Kammoona, J. H. Anderson, S. Brown, and T. Czajkowski, “Legup: high-level synthesis for fpga-based processor/accelerator systems,” in Proceedings of the 19th ACM/SIGDA international symposium on Field programmable gate arrays . ACM, 2011, pp. 33–36
T. Weyand, I. Kostrikov, and J. Philbin, “Planet-photo geolocation with convolutional neural networks,” in European Conference on Computer Vision . Springer, 2016, pp. 37–55
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Later among the works it cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: efficient inference engine on compressed deep neural network,” in Computer Architecture (ISCA), 2016 ACM/IEEE 43rd Annual International Symposium on . IEEE, 2016, pp. 243–254
2016
Later among the works it cites.
Y.-H. Chen, J. Emer, and V. Sze, “Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,” in ACM SIGARCH Computer Architecture News , vol. 44, no. 3. IEEE Press, 2016, pp. 367–379
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
S. W. Keckler, W. J. Dally, B. Khailany, M. Garland, and D. Glasco, “GPUs and the future of parallel computing,” IEEE Micro , vol. 31, no. 5, pp. 7–17, 2011
2011
Cited alongside, same era.
C. Farabet, Y. LeCun, K. Kavukcuoglu, E. Culurciello, B. Martini, P. Akselrod, and S. Talay, “Large-scale FPGA-based convolutional networks,” Scaling up Machine Learning: Parallel and Distributed Approaches , pp. 399–419, 2011
2011
Cited alongside, same era.
C. Farabet, B. Martini, B. Corda, P. Akselrod, E. Culurciello, and Y. LeCun, “Neuflow: A runtime reconfigurable dataflow processor for vision,” in Computer Vision and Pattern Recognition Workshops (CVPRW), 2011 IEEE Computer Society Conference on . IEEE, 2011, pp. 109–116
2011
Cited alongside, same era.
K. O. W. Group et al. , “The opencl specification version 1.1,” http://www. khronos. org/registry/cl/specs/opencl-1.1. pdf , 2011
2011
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Cited alongside, same era.
A.-r. Mohamed, G. E. Dahl, G. Hinton et al. , “Acoustic modeling using deep belief networks,” IEEE Trans. Audio, Speech & Language Processing , vol. 20, no. 1, pp. 14–22, 2012
2012
Cited alongside, same era.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal processing magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Cited alongside, same era.
H. Esmaeilzadeh, A. Sampson, L. Ceze, and D. Burger, “Neural acceleration for general-purpose approximate programs,” in Proceedings of the 2012 45th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Computer Society, 2012, pp. 449–460
2012
Cited alongside, same era.
Y. Ma, N. Suda, Y. Cao, J.-s. Seo, and S. Vrudhula, “Scalable and modularized rtl compilation of convolutional neural networks onto FPGA,” in Field Programmable Logic and Applications (FPL), 2016 26th International Conference on . IEEE, 2016, pp. 1–8
2016
Later among the works it cites.
N. Suda, V. Chandra, G. Dasika, A. Mohanty, Y. Ma, S. Vrudhula, J.-s. Seo, and Y. Cao, “Throughput-optimized opencl-based FPGA accelerator for large-scale convolutional neural networks,” in Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2016, pp. 16–25
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European conference on computer vision . Springer, 2016, pp. 630–645
2016
Later among the works it cites.
B. S. C. Varma, K. Paul, and M. Balakrishnan, Architecture exploration of FPGA based accelerators for BioInformatics applications . Springer, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
J. Qiu, J. Wang, S. Yao, K. Guo, B. Li, E. Zhou, J. Yu, T. Tang, N. Xu, and S. Song, “Going deeper with embedded FPGA platform for convolutional neural network,” in Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2016, pp. 26–35
2016
Later among the works it cites.
A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V. Srikumar, “Isaac: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars,” ACM SIGARCH Computer Architecture News , vol. 44, no. 3, pp. 14–26, 2016
2016
Later among the works it cites.
P. Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “Prime: A novel processing-in-memory architecture for neural network computation in reram-based main memory,” in ACM SIGARCH Computer Architecture News , vol. 44, no. 3. IEEE Press, 2016, pp. 27–39
2016
Later among the works it cites.
Y. Wang, J. Xu, Y. Han, H. Li, and X. Li, “Deepburning: automatic generation of FPGA-based learning accelerators for the neural network family,” in Proceedings of the 53rd Annual Design Automation Conference . ACM, 2016, p. 110
2016
Later among the works it cites.
C. Zhang, Z. Fang, P. Zhou, P. Pan, and J. Cong, “Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks,” in Computer-Aided Design (ICCAD), 2016 IEEE/ACM International Conference on . IEEE, 2016, pp. 1–8
2016
Later among the works it cites.
S. I. Venieris and C.-S. Bouganis, “FPGAconvnet: A framework for mapping convolutional neural networks on fpgas,” in Field-Programmable Custom Computing Machines (FCCM), 2016 IEEE 24th Annual International Symposium on . IEEE, 2016, pp. 40–47
2016
Later among the works it cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “Tensorflow: a system for large-scale machine learning.” in OSDI , vol. 16, 2016, pp. 265–283
2016
Later among the works it cites.
M. H. Alsuwaiyel, Algorithms: Design Techniques And Analysis (Revised Edition) . World Scientific, 2016, vol. 14
2016
Later among the works it cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “Xnor-net: Imagenet classification using binary convolutional neural networks,” in European Conference on Computer Vision . Springer, 2016, pp. 525–542
2016
Later among the works it cites.
M. Kim and P. Smaragdis, “Bitwise neural networks,” arXiv preprint arXiv:1601.06071 , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
M. Courbariaux and Y. Bengio, “Binarynet: Training deep neural networks with weights and activations constrained to+ 1 or- 1,” 2016
2016
Later among the works it cites.
A. Lavin and S. Gray, “Fast algorithms for convolutional neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 4013–4021
2016
Later among the works it cites.
2016
Later among the works it cites.
H. Li, X. Fan, L. Jiao, W. Cao, X. Zhou, and L. Wang, “A high performance FPGA-based accelerator for large-scale convolutional neural networks,” in Field Programmable Logic and Applications (FPL), 2016 26th International Conference on . IEEE, 2016, pp. 1–9
2016
Later among the works it cites.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in European conference on computer vision . Springer, 2016, pp. 21–37
2016
Later among the works it cites.
C. Zhang, D. Wu, J. Sun, G. Sun, G. Luo, and J. Cong, “Energy-efficient cnn implementation on a deeply pipelined fpga cluster,” in Proceedings of the 2016 International Symposium on Low Power Electronics and Design . ACM, 2016, pp. 326–331
2016
Later among the works it cites.
B. Wu, F. N. Iandola, P. H. Jin, and K. Keutzer, “Squeezedet: Unified, small, low power fully convolutional neural networks for real-time object detection for autonomous driving.” in CVPR Workshops , 2017, pp. 446–454
2017
Later among the works it cites.
A. Vasudevan, A. Anderson, and D. Gregg, “Parallel multi channel convolution using general matrix multiplication,” in Application-specific Systems, Architectures and Processors (ASAP), 2017 IEEE 28th International Conference on . IEEE, 2017, pp. 19–24
2017
Later among the works it cites.
E. Nurvitadhi, G. Venkatesh, J. Sim, D. Marr, R. Huang, J. Ong Gee Hock, Y. T. Liew, K. Srivatsan, D. Moss, S. Subhaschandra et al. , “Can FPGAs beat GPUs in accelerating next-generation deep neural networks?” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 5–14
2017
Later among the works it cites.
Y. Ma, Y. Cao, S. Vrudhula, and J.-s. Seo, “Optimizing loop operation and dataflow in FPGA acceleration of deep convolutional neural networks,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 45–54
2017
Later among the works it cites.
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, “Inception-v4, inception-resnet and the impact of residual connections on learning.” in AAAI , vol. 4, 2017, p. 12
2017
Later among the works it cites.
S. Han, J. Kang, H. Mao, Y. Hu, X. Li, Y. Li, D. Xie, H. Luo, S. Yao, Y. Wang et al. , “Ese: Efficient speech recognition engine with sparse lstm on FPGA,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 75–84
2017
Later among the works it cites.
T. Luo, S. Liu, L. Li, Y. Wang, S. Zhang, T. Chen, Z. Xu, O. Temam, and Y. Chen, “Dadiannao: A neural network supercomputer,” IEEE Transactions on Computers , vol. 66, no. 1, pp. 73–88, 2017
2017
Later among the works it cites.
V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proceedings of the IEEE , vol. 105, no. 12, pp. 2295–2329, 2017
2017
Later among the works it cites.
W. Lu, G. Yan, J. Li, S. Gong, Y. Han, and X. Li, “Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks,” in High Performance Computer Architecture (HPCA), 2017 IEEE International Symposium on . IEEE, 2017, pp. 553–564
2017
Later among the works it cites.
Z. Liu, Y. Dou, J. Jiang, J. Xu, S. Li, Y. Zhou, and Y. Xu, “Throughput-optimized fpga accelerator for deep convolutional neural networks,” ACM Transactions on Reconfigurable Technology and Systems (TRETS) , vol. 10, no. 3, p. 17, 2017
2017
Later among the works it cites.
Y. Guan, H. Liang, N. Xu, W. Wang, S. Shi, X. Chen, G. Sun, W. Zhang, and J. Cong, “FP-DNN: An automated framework for mapping deep neural networks onto FPGAs with RTL-HLS hybrid templates,” in Field-Programmable Custom Computing Machines (FCCM), 2017 IEEE 25th Annual International Symposium on . IEEE, 2017, pp. 152–159
2017
Later among the works it cites.
Y. Umuroglu, N. J. Fraser, G. Gambardella, M. Blott, P. Leong, M. Jahre, and K. Vissers, “Finn: A framework for fast, scalable binarized neural network inference,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 65–74
2017
Later among the works it cites.
S. I. Venieris and C.-S. Bouganis, “Latency-driven design for fpga-based convolutional neural networks,” in Field Programmable Logic and Applications (FPL), 2017 27th International Conference on . IEEE, 2017, pp. 1–8
2017
Later among the works it cites.
C. Zhang and V. Prasanna, “Frequency domain acceleration of convolutional neural networks on cpu-fpga shared memory system,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 35–44
2017
Later among the works it cites.
U. Aydonat, S. O’Connell, D. Capalija, A. C. Ling, and G. R. Chiu, “An opencl™ deep learning accelerator on arria 10,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 55–64
2017
Later among the works it cites.
L. Lu, Y. Liang, Q. Xiao, and S. Yan, “Evaluating fast algorithms for convolutional neural networks on fpgas,” in Field-Programmable Custom Computing Machines (FCCM), 2017 IEEE 25th Annual International Symposium on . IEEE, 2017, pp. 101–108
2017
Later among the works it cites.
J. Zhang and J. Li, “Improving the performance of opencl-based fpga accelerator for convolutional neural network,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 25–34
2017
Later among the works it cites.
Y. Shen, M. Ferdman, and P. Milder, “Maximizing cnn accelerator efficiency through resource partitioning,” in Computer Architecture (ISCA), 2017 ACM/IEEE 44th Annual International Symposium on . IEEE, 2017, pp. 535–547
2017
Later among the works it cites.
W. Xuechao, Y. Cody Hao, Z. Peng, C. Youxiang, W. Yuxin, H. Han, L. Yun, and C. Jason, “Automated systolic array architecture synthesis for high throughput cnn inference on FPGAs,” in Proceedings of the 2017 Design Automation Conference . ACM, 2017, pp. 1–6
2017
Later among the works it cites.
Y. Ma, M. Kim, Y. Cao, S. Vrudhula, and J.-s. Seo, “End-to-end scalable FPGA accelerator for deep residual networks,” in Circuits and Systems (ISCAS), 2017 IEEE International Symposium on . IEEE, 2017, pp. 1–4
2017
Later among the works it cites.
C. Wang, L. Gong, Q. Yu, X. Li, Y. Xie, and X. Zhou, “Dlau: A scalable deep learning accelerator unit on FPGA,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 36, no. 3, pp. 513–517, 2017
2017
Later among the works it cites.
Y. Ma, Y. Cao, S. Vrudhula, and J.-s. Seo, “An automatic rtl compiler for high-throughput FPGA implementation of diverse deep convolutional neural networks,” in Field Programmable Logic and Applications (FPL), 2017 27th International Conference on . IEEE, 2017, pp. 1–8
2017
Later among the works it cites.
E. Chung, J. Fowers, K. Ovtcharov, M. Papamichael, A. Caulfield, T. Massengil, M. Liu, D. Lo, S. Alkalay, M. Haselman et al. , “Accelerating persistent neural networks at datacenter scale,” in Hot Chips , vol. 27, 2017
2017
Later among the works it cites.
L. Xie and A. L. Yuille, “Genetic CNN.” in ICCV , 2017, pp. 1388–1397
2017
Later among the works it cites.
L. Zhang, S. Wang, and B. Liu, “Deep learning for sentiment analysis: A survey,” Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery , p. e1253, 2018
2018
Later among the works it cites.
Mathworks. What Is Deep Learning? [Online]. Available: https://www.mathworks.com/discovery/deep-learning.html/ , 2018
2018
Later among the works it cites.
A. Deshpande. A Beginner’s Guide To Understanding Convolutional Neural Networks [Online]. Available: https://adeshpande3.github.io/A-Beginner%27s-Guide-To-Understanding-Convolutional-Neural-Networks/ , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Image-Net. The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) [Online]. Available: http://image-net.org/challenges/LSVRC/ , 2018
2018
Later among the works it cites.
K. Guo, L. Sui, J. Qiu, J. Yu, J. Wang, S. Yao, S. Han, Y. Wang, and H. Yang, “Angel-eye: A complete design flow for mapping cnn onto embedded FPGA,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 37, no. 1, pp. 35–47, 2018
2018
Later among the works it cites.
L. Du, Y. Du, Y. Li, J. Su, Y.-C. Kuan, C.-C. Liu, and M.-C. F. Chang, “A reconfigurable streaming deep convolutional neural network accelerator for internet of things,” IEEE Transactions on Circuits and Systems I: Regular Papers , vol. 65, no. 1, pp. 198–208, 2018
2018
Later among the works it cites.
P. Joshi. What Is Local Response Normalization In Convolutional Neural Networks [Online]. Available: https://prateekvjoshi.com/2016/04/05/what-is-local-response-normalization-in-convolutional-neural-networks/ , 2018
2018
Later among the works it cites.
A. Karpathy. Convolutional Neural Networks for Visual Recognition [Online]. Available: http://cs231n.github.io/convolutional-networks/ , 2018
2018
Later among the works it cites.
H. M. Waidyasooriya, M. Hariyama, and K. Uchiyama, Design of FPGA-Based Computing Systems with OpenCL . Springer, 2018
2018
Later among the works it cites.
V. Sze, Y.-H. Chen, J. Emer, A. Suleiman, and Z. Zhang, “Hardware for machine learning: Challenges and opportunities,” in Custom Integrated Circuits Conference (CICC), 2018 IEEE . IEEE, 2018, pp. 1–8
2018
Later among the works it cites.
C. Zhang, G. Sun, Z. Fang, P. Zhou, P. Pan, and J. Cong, “Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , 2018
2018
Later among the works it cites.
Altera. OpenCL Design Examples [Online]. Available: https://www.altera.com/support/support-resources/designexamples/design-software/opencl.html/ , 2018
2018
Later among the works it cites.
Nallatech. P395-D8 OpenCL FPGA Accelerator Cards [Online]. Available: http://www.nallatech.com/wp-content/uploads/openclcardspb_v1_51.pdf/ , 2018
2018
Later among the works it cites.
Altera. DE5-Net FPGA Kit User Manual [Online]. Available: ftp://ftp.altera.com/up/pub/Altera_Material/Boards/DE5/DE5_User_/ , 2018
2018
Later among the works it cites.
Y. Ma, N. Suda, Y. Cao, S. Vrudhula, and J.-s. Seo, “Alamo: FPGA acceleration of deep learning algorithms with a modularized rtl compiler,” Integration , 2018
2018
Later among the works it cites.
Altera. JTAG UART Core [Online]. Available: https://www.altera.com/en_US/pdfs/literature/hb/nios2/n2cpu_nii51009.pdf , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. Ma, Y. Cao, S. Vrudhula, and J.-s. Seo, “Optimizing the convolution operation to accelerate deep neural networks on FPGA,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , no. 99, pp. 1–14, 2018
2018
Later among the works it cites.