Fetching the paper…
Reading the bibliography…
High-performance computing systems are moving towards 2.5D and 3D memory hierarchies, based on High Bandwidth Memory (HBM) and Hybrid Memory Cube (HMC) to mitigate the main memory bottlenecks.
S. Williams, A. Waterman, and D. Patterson, “Roofline: An insightful visual performance model for multicore architectures,” Commun. ACM , vol. 52, no. 4, pp. 65–76, Apr. 2009
2009
Earlier work this paper cites.
C. Farabet, B. Martini, B. Corda et al. , “NeuFlow: A runtime reconfigurable dataflow processor for vision,” in CVPR 2011 WORKSHOPS , June 2011, pp. 109–116
2011
Earlier work this paper cites.
P. H. Pham, D. Jelaca, C. Farabet et al. , “NeuFlow: Dataflow vision processing system-on-a-chip,” in IEEE 55th International Midwest Symposium on Circuits and Systems (MWSCAS) , 2012, pp. 1044–1047
2012
Earlier work this paper cites.
J. Jeddeloh and B. Keeth, “Hybrid memory cube new DRAM architecture increases density and performance,” in VLSI Technology (VLSIT), 2012 Symposium on , June 2012, pp. 87–88
2012
Earlier work this paper cites.
S. E. Kahou, C. Pal, X. Bouthillier et al. , “Combining modality specific deep neural networks for emotion recognition in video,” in Proceedings of the 15th ACM on International Conference on Multimodal Interaction , ser. ICMI ’13, 2013, pp. 543–550
2013
Earlier work this paper cites.
S. Carrillo, J. Harkin, L. J. McDaid et al. , “Scalable hierarchical network-on-chip architecture for spiking neural network hardware implementations,” IEEE Transactions on Parallel and Distributed Systems , vol. 24, no. 12, pp. 2451–2461, Dec 2013
2013
Earlier work this paper cites.
A. Haidar, J. Kurzak, and P. Luszczek, “An improved parallel singular value algorithm and its implementation for multicore hardware,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
G. Kim, J. Kim, J. H. Ahn, and J. Kim, “Memory-centric system interconnect design with hybrid memory cubes,” in Proceedings of the 22Nd International Conference on Parallel Architectures and Compilation Techniques , ser. PACT ’13, 2013, pp. 145–156
2013
Earlier work this paper cites.
E. Azarkhish, I. Loi, and L. Benini, “A case for three-dimensional stacking of tightly coupled data memories over multi-core clusters using low-latency interconnects,” IET Computers Digital Techniques , vol. 7, no. 5, pp. 191–199, September 2013
2013
Earlier work this paper cites.
P. M. Kogge and D. R. Resnick, “Yearly update: Exascale projections for 2013,” Sandia National Laboratoris, Tech. Rep. SAND2013-9229, Oct. 2013
2013
Earlier work this paper cites.
Y. Taigman, M. Yang, M. Ranzato, and L. Wolf, “DeepFace: Closing the gap to human-level performance in face verification,” in Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie et al. , Microsoft COCO: Common Objects in Context . Springer International Publishing, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue et al. , “Caffe: Convolutional architecture for fast feature embedding,” in Proceedings of the 22nd ACM International Conference on Multimedia , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
H. Kimura, P. Aziz, T. Jing et al. , “28Gb/s 560mW multi-standard SerDes with single-stage analog front-end and 14-tap decision-feedback equalizer in 28nm CMOS,” in 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) , Feb 2014
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia et al. , “Going deeper with convolutions,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2015, pp. 1–9
2015
Earlier work this paper cites.
J. Bilski and J. Smolag, “Parallel architectures for learning the RTRN and Elman dynamic neural networks,” IEEE Transactions on Parallel and Distributed Systems , vol. 26, no. 9, pp. 2561–2570, Sept 2015
2015
Earlier work this paper cites.
C. Zhang, P. Li, G. Sun et al. , “Optimizing FPGA-based accelerator design for deep convolutional neural networks,” in Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
2015
Cited alongside, same era.
T. Chen, Z. Du, N. Sun et al. , “A high-throughput neural network accelerator,” IEEE Micro , vol. 35, no. 3, pp. 24–32, May 2015
2015
Cited alongside, same era.
Z. Du, R. Fasthuber, T. Chen et al. , “ShiDianNao: Shifting vision processing closer to the sensor,” SIGARCH Comput. Archit. News , vol. 43, no. 3, pp. 92–104, Jun. 2015
2015
Cited alongside, same era.
L. Xu, D. Zhang, and N. Jayasena, “Scaling deep learning on multiple in-memory processors,” in 3rd Workshop on Near-Data Processing (WoNDP) , 2015
2015
Cited alongside, same era.
Hybrid Memory Cube Specification 2.1 , Hybrid Memory Cube Consortium Std., 2015
E. Azarkhish, C. Pfister, D. Rossi, I. Loi, and L. Benini, “Logic-base interconnect design for near memory computing in the Smart Memory Cube,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. PP, no. 99, pp. 1–14, 2016
2016
Later among the works it cites.
D. Kang, W. Jeong, C. Kim et al. , “256Gb 3b/cell V-NAND flash memory with 48 stacked WL layers,” in 2016 IEEE International Solid-State Circuits Conference (ISSCC) , Jan 2016, pp. 130–131
2016
Later among the works it cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016, http://www.deeplearningbook.org
2016
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
L. Cavigelli, M. Magno, and L. Benini, “Accelerating real-time embedded scene labeling with convolutional networks,” in 2015 52nd ACM/EDAC/IEEE Design Automation Conference (DAC) , 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
I. Loi, D. Rossi et al. , “Exploring multi-banked shared-L1 program cache on ultra-low power, tightly coupled processor clusters,” in Proceedings of the 12th ACM International Conference on Computing Frontiers , 2015
2015
Cited alongside, same era.
E. Azarkhish, D. Rossi, I. Loi, and L. Benini, “A modular shared L2 memory design for 3-D integration,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 23, no. 8, pp. 1485–1498, 2015
2015
Cited alongside, same era.
S. Srinivas, R. K. Sarvadevabhatla, K. R. Mopuri et al. , “A taxonomy of deep convolutional neural nets for computer vision,” Frontiers in Robotics and AI , vol. 2, p. 36, 2016
2016
Cited alongside, same era.
X. Li, Y. Zhang, M. Li et al. , “Deep neural network for RFID-based activity recognition,” in Proceedings of the Eighth Wireless of the Students, by the Students, and for the Students Workshop , ser. S3, 2016, pp. 24–26
2016
Cited alongside, same era.
2016
Cited alongside, same era.
M. Peemen, R. Shi, S. Lal et al. , “The neuro vector engine: Flexibility to improve convolutional net efficiency for wearable vision,” in 2016 Design, Automation Test in Europe Conference Exhibition (DATE) , March 2016, pp. 1604–1609
2016
Later among the works it cites.
J. Gómez-Luna, I. J. Sung, L. W. Chang et al. , “In-place matrix transposition on GPUs,” IEEE Transactions on Parallel and Distributed Systems , vol. 27, no. 3, pp. 776–788, March 2016
2016
Later among the works it cites.
S. Liu, Z. Du, J. Tao et al. , “Cambricon: An instruction set architecture for neural network,” in Proceedings of the 43rd Annual International Symposium on Computer Architecture , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
X. Meng, J. Bradley, B. Yavuz et al. , “MLlib: Machine learning in Apache Spark,” J. Mach. Learn. Res. , vol. 17, no. 1, pp. 1235–1241, Jan. 2016
2016
Later among the works it cites.
M. Grossman, M. Breternitz, and V. Sarkar, “HadoopCL2: Motivating the design of a distributed, heterogeneous programming system with machine-learning applications,” IEEE Transactions on Parallel and Distributed Systems , vol. 27, no. 3, pp. 762–775, March 2016
2016
Later among the works it cites.
S. Mittal and J. S. Vetter, “A survey of software techniques for using non-volatile memories for storage and main memory systems,” IEEE Transactions on Parallel and Distributed Systems , vol. 27, no. 5, pp. 1537–1550, May 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
D. Rossi, A. Pullini, I. Loi et al. , “A 60 GOPS/W, -1.8 V to 0.9 V body bias ULP cluster in 28nm UTBB FD-SOI technology,” Solid-State Electronics , vol. 117, pp. 170 – 184, 2016
2016
Later among the works it cites.
K. Sohn, W. J. Yun, R. Oh et al. , “A 1.2V 20nm 307GB/s HBM DRAM with at-speed wafer-level I/O test scheme and adaptive refresh considering temperature distribution,” in 2016 IEEE International Solid-State Circuits Conference (ISSCC) , Jan 2016, pp. 316–317
2016
Later among the works it cites.
M. Gao, J. Pu, X. Yang, M. Horowitz, and C. Kozyrakis, “Tetris: Scalable and efficient neural network acceleration with 3d memory,” SIGARCH Comput. Archit. News , vol. 45, no. 1, pp. 751–764, Apr. 2017
2017
Closest in time.
Y. H. Chen, T. Krishna, J. S. Emer, and V. Sze, “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE Journal of Solid-State Circuits , vol. 52, no. 1, pp. 127–138, Jan 2017
2017
Closest in time.
“Deep learning inference platform performance study,” White Paper, NVIDIA, 2017
2017
Closest in time.