Fetching the paper…
Reading the bibliography…
The data partitioning and scheduling strategies used by DNN accelerators to leverage reuse and perform staging are known as dataflow, and they directly impact the performance and energy efficiency of DNN accelerator designs.
Learning hierarchical features for scene labeling
Clement Farabet, Camille Couprie, Laurent Najman, and Yann LeCun. 2013 · 1929
Earlier work this paper cites.
A Data Locality Optimizing Algorithm. In Proceedings of the ACM SIGPLAN 1991 Conference on Programming Language Design and Implementation (PLDI ’91) . ACM, New York, NY, USA, 30–44
Michael E. Wolf and Monica S. Lam. 1991 · 1991
Earlier work this paper cites.
Data-centric Multi-level Blocking. In Proceedings of the ACM SIGPLAN 1997 Conference on Programming Language Design and Implementation (PLDI ’97) . ACM, New York, NY, USA, 346–357
Induprakas Kodukula, Nawaaz Ahmed, and Keshav Pingali. 1997 · 1997
Earlier work this paper cites.
Automatic selection of high-order transformations in the IBM XL FORTRAN compilers
Vivek Sarkar. 1997 · 1997
Earlier work this paper cites.
An Experimental Evaluation of Tiling and Shackling for Memory Hierarchy Management. In Proceedings of the 13th International Conference on Supercomputing (ICS ’99) . ACM, New York, NY, USA, 482–491
Induprakas Kodukula, Keshav Pingali, Robert Cox, and Dror Maydan. 1999 · 1999
Earlier work this paper cites.
An analytical model for loop tiling and its solution. In ISPASS . 146–153
Vivek Sarkar and Nimrod Megiddo. 2000 · 2000
Earlier work this paper cites.
Data-Centric Transformations for Locality Enhancement
Induprakas Kodukula and Keshav Pingali. 2001 · 2001
Earlier work this paper cites.
A Practical Automatic Polyhedral Parallelizer and Locality Optimizer. In Proceedings of the 29th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI ’08) . ACM, New York, NY, USA, 101–113
Uday Bondhugula, Albert Hartono, J. Ramanujam, and P. Sadayappan. 2008 · 2008
Earlier work this paper cites.
CACTI 6.0: A tool to model large caches
Naveen Muralimanohar, Rajeev Balasubramonian, and Norman P Jouppi. 2009 · 2009
Earlier work this paper cites.
Combined Iterative and Model-driven Optimization in an Automatic Parallelization Framework. In Conference on High Performance Computing Networking, Storage and Analysis, SC 2010, New Orleans, LA, USA, November 13-19, 2010 . 1–11
Louis-Noël Pouchet, Uday Bondhugula, Cédric Bastoul, Albert Cohen, J. Ramanujam, and P. Sadayappan. 2010 · 2010
Earlier work this paper cites.
Analytical Bounds for Optimal Tile Size Selection. In Compiler Construction - 21st International Conference, CC 2012, Held as Part of the European Joint Conferences on Theory and Practice of Software, ETAPS 2012, Tallinn, Estonia, March 24 - April 1, 2012. Proceedings . 101–121
Jun Shirako, Kamal Sharma, Naznin Fauzia, Louis-Noël Pouchet, J. Ramanujam, P. Sadayappan, and Vivek Sarkar. 2012 · 2012
Earlier work this paper cites.
Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning. In International conference on Architectural support for programming languages and operating systems (ASPLOS)
Tianshi Chen, Zidong Du, Ninghui Sun, Jia Wang, Chengyong Wu, Yunji Chen, and Olivier Temam. 2014 · 2014
Earlier work this paper cites.
Minimizing computation in convolutional neural networks. In International conference on artificial neural networks (ICANN) . Springer, 281–290
Jason Cong and Bingjun Xiao. 2014 · 2014
Earlier work this paper cites.
Oil and Water Can Mix: An Integration of Polyhedral and AST-based Transformations. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’14) . IEEE Press, Piscataway, NJ, USA, 287–298
Jun Shirako, Louis-Noël Pouchet, and Vivek Sarkar. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Deeppose: Human pose estimation via deep neural networks. In Conference on Computer Vision and Pattern Recognition (CVPR)
Alexander Toshev and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
Luke Metz Alec Radford and Soumith Chintala. 2015 · 2015
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, et al · 2015
Cited alongside, same era.
ShiDianNao: Shifting vision processing closer to the sensor. In International Symposium on Computer Architecture (ISCA)
Zidong Du, Robert Fasthuber, Tianshi Chen, Paolo Ienne, Ling Li, Tao Luo, Xiaobing Feng, Yunji Chen, and Olivier Temam. 2015 · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions. In Conference on Computer Vision and Pattern Recognition (CVPR)
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Tetris: Scalable and efficient neural network acceleration with 3d memory
Mingyu Gao, Jing Pu, Xuan Yang, Mark Horowitz, and Christos Kozyrakis. 2017 · 2017
Later among the works it cites.
In-datacenter performance analysis of a tensor processing unit. In International Symposium on Computer Architecture (ISCA) . IEEE, 1–12
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Later among the works it cites.
FlexFlow: A Flexible Dataflow Accelerator Architecture for Convolutional Neural Networks. In International Symposium on High Performance Computer Architecture (HPCA)
Wenyan Lu, Guihai Yan, Jiajun Li, Shijun Gong, Yinhe Han, and Xiaowei Li. 2017 · 2017
Later among the works it cites.
Optimizing loop operation and dataflow in fpga acceleration of deep convolutional neural networks. In International Symposium on Field-Programmable Gate Arrays (FPGA) . ACM, 45–54
Yufei Ma, Yu Cao, Sarma Vrudhula, and Jae-sun Seo. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention . Springer, 234–241
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Cited alongside, same era.
Very Deep Convolutional Networks For Large-Scale Image Recognition. In International Conference on Learning Representations (ICLR)
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks. In International Symposium on Computer Architecture (ISCA)
Yu-Hsin Chen, Joel Emer, and Vivienne Sze. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
From high-level deep neural models to FPGAs. In IEEE/ACM International Symposium on Microarchitecture (MICRO)
Hardik Sharma, Jongse Park, Divya Mahajan, Emmanuel Amaro, Joon Kyung Kim, Chenkai Shao, Asit Mishra, and Hadi Esmaeilzadeh. 2016 · 2016
Cited alongside, same era.
C-brain: a deep learning accelerator that tames the diversity of CNNs through adaptive data-level parallelization. In Design Automation Conference (DAC) . 1–6
Lili Song, Ying Wang, Yinhe Han, Xin Zhao, Bosheng Liu, and Xiaowei Li. 2016 · 2016
Cited alongside, same era.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Cited alongside, same era.
NVDLA Deep Learning Accelerator
2017 · 2017
Cited alongside, same era.
SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks. In International Symposium on Computer Architecture (ISCA) . 27–40
A. Parashar et al · 2017
Later among the works it cites.
Polyhedral Optimization of TensorFlow Computation Graphs. In Workshop on Extreme-scale Programming Tools (ESPT)
Benoît Pradelle, Benoît Meister, M. Baskaran, Jonathan Springer, and Richard Lethin. 2017 · 2017
Later among the works it cites.
Aggregated Residual Transformations for Deep Neural Networks
Piotr Dollár Zhuowen Tu Saining Xie, Ross Girshick and Kaiming He. 2017 · 2017
Later among the works it cites.
Snapea: Predictive early activation for reducing computation in deep convolutional neural networks. In International Symposium on Computer Architecture (ISCA)
V Aklaghi, Amir Yazdanbakhsh, Kambiz Samadi, H Esmaeilzadeh, and RK Gupta. 2018 · 2018
Closest in time.
MAERI: Enabling Flexible Dataflow Mapping over DNN Accelerators via Reconfigurable Interconnects. In International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) . 461–475
Hyoukjun Kwon, Ananda Samajdar, and Tushar Krishna. 2018 · 2018
Closest in time.
Diffy: a Déja vu-Free Differential Deep Neural Network Accelerator. In International Symposium on Microarchitecture (MICRO)
Mostafa Mahmoud, Kevin Siu, and Andreas Moshovos. 2018 · 2018
Closest in time.
Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks
Chen Zhang, Guangyu Sun, Zhenman Fang, Peipei Zhou, Peichen Pan, and Jason Cong. 2018 · 2018
Closest in time.
MAESTRO project page
2019 · 2019
Closest in time.
MobileNetV2: Inverted Residuals and Linear Bottlenecks
Menglong Zhu Andrey Zhmoginov Mark Sandler, Andrew Howard and Liang-Chieh Chen. 2019 · 2019
Closest in time.
Timeloop: A Systematic Approach to DNN Accelerator Evaluation. In Proceedings of the 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)
Angshuman Parashar, Priyanka Raina, Yakun Sophia Shao, Yu-Hsin Chen, Victor A. Ying, Anurag Mukkara, Rangharajan Venkatesan, Brucek Khailany, Stephen W. Keckler, and Joel Emer. 2019 · 2019
Closest in time.
Accelergy: An Architecture-Level Energy Estimation Methodology for Accelerator Designs. In IEEE/ACM International Conference On Computer Aided Design (ICCAD)
Wu, Yannan N. and Emer, Joel S. and Sze, Vivienne. 2019 · 2019
Closest in time.