Fetching the paper…
Reading the bibliography…
Machine learning (ML) models are widely used in many important domains.
I. S. Duff, “A survey of sparse matrix research,” Proceedings of the IEEE , vol. 65, no. 4, pp. 500–535, 1977
1977
Earlier work this paper cites.
D. R. Kincaid and D. M. Young, “The itpack project: Past, present, and future,” in Elliptic Problem Solvers . Elsevier, 1984, pp. 53–63
1984
Earlier work this paper cites.
Y. Saad, “Sparskit: A basic tool kit for sparse matrix computations,” 1990
1990
Earlier work this paper cites.
P. Feautrier, “Dataflow analysis of array and scalar references,” International Journal of Parallel Programming , vol. 20, no. 1, 1991
1991
Earlier work this paper cites.
M. E. Wolf, “Improving locality and parallelism in nested loops,” Ph.D. dissertation, to the Department of Computer Science.Stanford University, 1992
1992
Earlier work this paper cites.
E.-J. Im and K. Yelick, “Model-based memory hierarchy optimizations for sparse matrices,” in Workshop on Profile and Feedback-Directed Compilation , vol. 139, 1998
1998
Earlier work this paper cites.
E. Jones, T. Oliphant, and P. Peterson, “Scipy: Open source scientific tools for python,” 2001
2001
Earlier work this paper cites.
Y. Saad, Iterative methods for sparse linear systems . siam, 2003
2003
Earlier work this paper cites.
R. W. Vuduc and J. W. Demmel, Automatic performance tuning of sparse matrix kernels . University of California, Berkeley, 2003, vol. 1
2003
Earlier work this paper cites.
C. Lattner and V. Adve, “Llvm: A compilation framework for lifelong program analysis & transformation,” in International Symposium on Code Generation and Optimization, 2004. CGO 2004. , 2004
2004
Earlier work this paper cites.
J. Willcock and A. Lumsdaine, “Accelerating sparse matrix computations via data compression,” in Proceedings of the 20th annual international conference on Supercomputing , 2006, pp. 307–316
2006
Earlier work this paper cites.
F. Agakov, E. Bonilla, J. Cavazos, B. Franke, G. Fursin, M. F. O’Boyle, J. Thomson, M. Toussaint, and C. K. Williams, “Using machine learning to focus iterative optimization,” in International Symposium on Code Generation and Optimization (CGO’06) . IEEE, 2006
2006
Earlier work this paper cites.
A. Buluc and J. R. Gilbert, “On the representation and multiplication of hypersparse matrices,” in 2008 IEEE International Symposium on Parallel and Distributed Processing . IEEE, 2008, pp. 1–11
2008
Earlier work this paper cites.
B. W. Bader and T. G. Kolda, “Efficient matlab computations with sparse and factored tensors,” SIAM Journal on Scientific Computing , vol. 30, no. 1, pp. 205–231, 2008
2008
Earlier work this paper cites.
U. Bondhugula, A. Hartono, J. Ramanujam, and P. Sadayappan, “A practical automatic polyhedral parallelizer and locality optimizer,” in PLDI , 2008, pp. 101–113
2008
Earlier work this paper cites.
C. Chen, J. Chame, and M. Hall, “Chill: A framework for composing high-level loop transformations,” U. of Southern California, Tech. Rep. 08-897, 2008
2008
Earlier work this paper cites.
K. Goto and R. A. v. d. Geijn, “Anatomy of high-performance matrix multiplication,” ACM Transactions on Mathematical Software (TOMS) , vol. 34, no. 3, pp. 1–25, 2008
2008
Earlier work this paper cites.
A. Buluç, J. T. Fineman, M. Frigo, J. R. Gilbert, and C. E. Leiserson, “Parallel sparse matrix-vector and matrix-transpose-vector multiplication using compressed sparse blocks,” in Proceedings of the twenty-first annual symposium on Parallelism in algorithms and architectures , 2009, pp. 233–244
2009
Earlier work this paper cites.
A. Hartono, M. M. Baskaran, C. Bastoul, A. Cohen, S. Krishnamoorthy, B. Norris, J. Ramanujam, and P. Sadayappan, “Parametric multi-level tiling of imperfectly nested loops,” in Proceedings of the 23rd international conference on Supercomputing , 2009, pp. 147–157
2009
Earlier work this paper cites.
K. Trifunovic, D. Nuzman, A. Cohen, A. Zaks, and I. Rosen, “Polyhedral-model guided loop-nest auto-vectorization,” in 2009 18th International Conference on Parallel Architectures and Compilation Techniques . IEEE, 2009, pp. 327–337
2009
Earlier work this paper cites.
D. Vainbrand and R. Ginosar, “Network-on-chip architectures for neural networks,” in 2010 Fourth ACM/IEEE International Symposium on Networks-on-Chip . IEEE, 2010, pp. 135–144
2010
Earlier work this paper cites.
S. Verdoolaege, “isl: An integer set library for the polyhedral model,” in International Congress on Mathematical Software . Springer, 2010
2010
Earlier work this paper cites.
M.-W. Benabderrahmane, L.-N. Pouchet, A. Cohen, and C. Bastoul, “The polyhedral model is more widely applicable than you think,” in Proceedings of the 19th Joint European Conference on Theory and Practice of Software, International Conference on Compiler Construction , ser. CC’10/ETAPS’10. Springer-Verlag, 2010
2010
Earlier work this paper cites.
C.-C. Chang and C.-J. Lin, “Libsvm: A library for support vector machines,” ACM transactions on intelligent systems and technology (TIST) , vol. 2, no. 3, pp. 1–27, 2011
2011
Earlier work this paper cites.
Y. Kim, J. Lee, A. Shrivastava, J. W. Yoon, D. Cho, and Y. Paek, “High throughput data mapping for coarse-grained reconfigurable architectures,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 30, no. 11, pp. 1599–1609, 2011
2011
Earlier work this paper cites.
F. Paul and L. Christian, “The polyhedron model,” in Encyclopedia of Parallel Computing , D. Padua, Ed. Springer, 2011, pp. 1581, 1592
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
M. Baskaran, B. Meister, N. Vasilache, and R. Lethin, “Efficient and scalable computations with sparse tensors,” in 2012 IEEE Conference on High Performance Extreme Computing . IEEE, 2012, pp. 1–6
2012
Earlier work this paper cites.
J. Ragan-Kelley, A. Adams, S. Paris, M. Levoy, S. Amarasinghe, and F. Durand, “Decoupling algorithms from schedules for easy optimization of image processing pipelines,” ACM Transactions on Graphics (TOG) , vol. 31, no. 4, pp. 1–12, 2012
2012
Earlier work this paper cites.
T. Grosser, A. Groslinger, and C. Lengauer, “Polly - performing polyhedral optimizations on a low-level intermediate representation.” Parallel Processing Letters , vol. 22, no. 4, 2012. [Online]. Available: http://dblp.uni-trier.de/db/journals/ppl/ppl22.html#GrosserGL12
2012
Earlier work this paper cites.
T. Yuki, G. Gupta, D. Kim, T. Pathan, and S. Rajopadhye, “Alphaz: A system for design space exploration in the polyhedral model,” in International Workshop on Languages and Compilers for Parallel Computing . Springer, 2012, pp. 17–31
2012
Earlier work this paper cites.
J. Bachrach, H. Vo, B. Richards, Y. Lee, A. Waterman, R. Avižienis et al. , “Chisel: constructing hardware in a scala embedded language,” in DAC Design Automation Conference 2012 , 2012, pp. 1212–1221
2012
Earlier work this paper cites.
N. I. of Standards and Technology, “Matrix market exchange formats,” https://math.nist.gov/MatrixMarket/formats.html , 2013
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems , 2014
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research , 2014
2014
Earlier work this paper cites.
E. Wang, Q. Zhang, B. Shen, G. Zhang, X. Lu, Q. Wu, and Y. Wang, “Intel math kernel library,” in High-Performance Computing on the Intel® Xeon Phi™ . Springer, 2014, pp. 167–188
2014
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun et al. , “Dadiannao: A machine-learning supercomputer,” in Proceedings of the 47th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Computer Society, 2014, pp. 609–622
2014
Earlier work this paper cites.
M. Horowitz, “1.1 computing’s energy problem (and what we can do about it),” in 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) . IEEE, 2014, pp. 10–14
2014
Earlier work this paper cites.
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,” in ACM Sigplan Notices , vol. 49, no. 4, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in Proceedings of the 22nd ACM international conference on Multimedia . ACM, 2014, pp. 675–678
2014
Earlier work this paper cites.
L. Wu, A. Lottarini, T. K. Paine, M. A. Kim, and K. A. Ross, “Q100: the architecture and design of a database processing unit,” in Proceedings of the 19th international conference on Architectural support for programming languages and operating systems , 2014
2014
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proceedings of the IEEE international conference on computer vision , 2015
2015
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in Advances in neural information processing systems , 2015, pp. 1135–1143
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Cichocki, D. Mandic, L. De Lathauwer, G. Zhou, Q. Zhao, C. Caiafa, and H. A. Phan, “Tensor decompositions for signal processing applications: From two-way to multiway component analysis,” IEEE signal processing magazine , vol. 32, no. 2, pp. 145–163, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1–9
2015
Earlier work this paper cites.
P. Warden, “Why are eight bits enough for deep neural networks?” https://petewarden.com/2015/05/23/why-are-eight-bits-enough-for-deep-neural-networks/ , 2015
2015
Earlier work this paper cites.
S. Smith and G. Karypis, “Tensor-matrix products with a compressed sparse tensor,” in Proceedings of the 5th Workshop on Irregular Applications: Architectures and Algorithms , 2015, pp. 1–7
2015
Earlier work this paper cites.
B. W. Bader, T. G. Kolda et al. , “Matlab tensor toolbox version 2.6,” Available online, February 2015. [Online]. Available: http://www.sandia.gov/~tgkolda/TensorToolbox/
2015
Earlier work this paper cites.
N. H. Weste and D. Harris, CMOS VLSI design: a circuits and systems perspective . Pearson Education India, 2015
2015
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) . IEEE, 2015, pp. 167–174
2015
Earlier work this paper cites.
R. Baghdadi, U. Beaugnon, A. Cohen, T. Grosser, M. Kruse, C. Reddy, S. Verdoolaege, A. Betts, A. F. Donaldson, J. Ketema et al. , “Pencil: A platform-neutral compute intermediate language for accelerator programming,” in 2015 International Conference on Parallel Architecture and Compilation (PACT) . IEEE, 2015, pp. 138–149
2015
Earlier work this paper cites.
R. T. Mullapudi, V. Vasista, and U. Bondhugula, “Polymage: Automatic optimization for image processing pipelines,” in Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems , 2015, pp. 429–443
2015
Earlier work this paper cites.
R. Baghdadi, A. Cohen, T. Grosser, S. Verdoolaege, A. Lokhmotov, J. Absar, S. van Haastregt, A. Kravets, and A. F. Donaldson, “PENCIL language specification,” INRIA, Research Rep. RR-8706, 2015. [Online]. Available: https://hal.inria.fr/hal-01154812
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Ahn, S. Hong, S. Yoo, O. Mutlu, and K. Choi, “A scalable processing-in-memory accelerator for parallel graph processing,” in Proceedings of the 42nd Annual International Symposium on Computer Architecture , 2015, pp. 105–117
2015
Earlier work this paper cites.
J. Fowers, J.-Y. Kim, D. Burger, and S. Hauck, “A scalable high-bandwidth architecture for lossless compression on fpgas,” in 2015 IEEE 23rd Annual International Symposium on Field-Programmable Custom Computing Machines . IEEE, 2015, pp. 52–59
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
Y.-H. Chen, T. Krishna, J. S. Emer, and V. Sze, “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE Journal of Solid-State Circuits , vol. 52, no. 1, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Zhang, Z. Du, L. Zhang, H. Lan, S. Liu, L. Li, Q. Guo, T. Chen, and Y. Chen, “Cambricon-x: An accelerator for sparse neural networks,” in The 49th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Press, 2016, p. 20
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: efficient inference engine on compressed deep neural network,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2016, pp. 243–254
2016
Earlier work this paper cites.
N. Suda, V. Chandra, G. Dasika, A. Mohanty, Y. Ma, S. Vrudhula, J.-s. Seo, and Y. Cao, “Throughput-optimized opencl-based fpga accelerator for large-scale convolutional neural networks,” in Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2016, pp. 16–25
2016
Earlier work this paper cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” ACM SIGARCH Computer Architecture News , 2016
2016
Earlier work this paper cites.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. K. Lee, J. M. Hernández-Lobato, G.-Y. Wei, and D. Brooks, “Minerva: Enabling low-power, highly-accurate deep neural network accelerators,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2016, pp. 267–278
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016
2016
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “Squad: 100,000+ questions for machine comprehension of text,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , 2016, pp. 2383–2392
2016
Earlier work this paper cites.
P. A. Tew, “An investigation of sparse tensor formats for tensor libraries,” Ph.D. dissertation, Massachusetts Institute of Technology, 2016
2016
Earlier work this paper cites.
J. King, T. Gilray, R. M. Kirby, and M. Might, “Dynamic sparse-matrix allocation on gpus,” in International Conference on High Performance Computing . Springer, 2016, pp. 61–80
2016
Earlier work this paper cites.
M. Alwani, H. Chen, M. Ferdman, and P. Milder, “Fused-layer cnn accelerators,” in 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2016, pp. 1–12
2016
Earlier work this paper cites.
H. Sharma, J. Park, D. Mahajan, E. Amaro, J. K. Kim, C. Shao, A. Mishra, and H. Esmaeilzadeh, “From high-level deep neural models to fpgas,” in The 49th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Press, 2016, p. 17
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in International conference on machine learning , 2016, pp. 173–182
2016
Earlier work this paper cites.
L. Truong, R. Barik, E. Totoni, H. Liu, C. Markley, A. Fox, and T. Shpeisman, “Latte: a language, compiler, and runtime for elegant and efficient deep neural networks,” in Proceedings of the 37th ACM SIGPLAN Conference on Programming Language Design and Implementation , 2016, pp. 209–223
2016
Earlier work this paper cites.
T. J. Ham, L. Wu, N. Sundaram, N. Satish, and M. Martonosi, “Graphicionado: A high-performance and energy-efficient accelerator for graph analytics,” in 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2016, pp. 1–13
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017
2017
Earlier work this paper cites.
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. Van Der Laak, B. Van Ginneken, and C. I. Sánchez, “A survey on deep learning in medical image analysis,” Medical image analysis , vol. 42, pp. 60–88, 2017
2017
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa et al. , “In-datacenter performance analysis of a tensor processing unit,” in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2017, pp. 1–12
2017
Earlier work this paper cites.
A. K. Mishra, E. Nurvitadhi, G. Venkatesh, J. Pearce, and D. Marr, “Fine-grained accelerators for sparse machine learning workloads,” in 2017 22nd Asia and South Pacific Design Automation Conference (ASP-DAC) . IEEE, 2017, pp. 635–640
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T.-J. Yang, Y.-H. Chen, and V. Sze, “Designing energy-efficient convolutional neural networks using energy-aware pruning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5687–5695
2017
Earlier work this paper cites.
V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proceedings of the IEEE , vol. 105, no. 12, pp. 2295–2329, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Yu, A. Lukefahr, D. Palframan, G. Dasika, R. Das, and S. Mahlke, “Scalpel: Customizing dnn pruning to the underlying hardware parallelism,” ACM SIGARCH Computer Architecture News , 2017
2017
Earlier work this paper cites.
S. Han, J. Kang, H. Mao, Y. Hu, X. Li, Y. Li, D. Xie, H. Luo, S. Yao, Y. Wang et al. , “Ese: Efficient speech recognition engine with sparse lstm on fpga,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , 2017, pp. 75–84
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Engelcke, D. Rao, D. Z. Wang, C. H. Tong, and I. Posner, “Vote3deep: Fast object detection in 3d point clouds using efficient convolutional neural networks,” in 2017 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2017, pp. 1355–1361
2017
Cited alongside, same era.
X. He, L. Liao, H. Zhang, L. Nie, X. Hu, and T.-S. Chua, “Neural collaborative filtering,” in Proceedings of the 26th international conference on world wide web , 2017, pp. 173–182
2017
Cited alongside, same era.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017
2017
Cited alongside, same era.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Quantized neural networks: Training neural networks with low precision weights and activations,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 6869–6898, 2017
2017
2019
Later among the works it cites.
C. Zhang, P. Patras, and H. Haddadi, “Deep learning in mobile and wireless networking: A survey,” IEEE Communications Surveys & Tutorials , 2019
2019
Later among the works it cites.
A. Ren, T. Zhang, S. Ye, J. Li, W. Xu, X. Qian, X. Lin, and Y. Wang, “Admm-nn: An algorithm-hardware co-design framework of dnns using alternating direction methods of multipliers,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems , 2019, pp. 925–938
2019
Later among the works it cites.
Y.-H. Chen, T.-J. Yang, J. Emer, and V. Sze, “Eyeriss v2: A flexible accelerator for emerging deep neural networks on mobile devices,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems , vol. 9, no. 2, pp. 292–308, 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
E. H. Lee, D. Miyashita, E. Chai, B. Murmann, and S. S. Wong, “Lognet: Energy-efficient neural networks using logarithmic computation,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 5900–5904
2017
Cited alongside, same era.
J. Albericio, A. Delmás, P. Judd, S. Sharify, G. O’Leary, R. Genov, and A. Moshovos, “Bit-pragmatic deep neural network computing,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture , 2017, pp. 382–394
2017
Cited alongside, same era.
B. Moons, R. Uytterhoeven, W. Dehaene, and M. Verhelst, “14.5 envision: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy-frequency-scalable convolutional neural network processor in 28nm fdsoi,” in 2017 IEEE International Solid-State Circuits Conference (ISSCC) . IEEE, 2017, pp. 246–247
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,” in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2017, pp. 27–40
2017
Cited alongside, same era.
L. Yavits and R. Ginosar, “Accelerator for sparse machine learning,” IEEE Computer Architecture Letters , vol. 17, no. 1, pp. 21–24, 2017
2017
Cited alongside, same era.
A. Page, A. Jafari, C. Shea, and T. Mohsenin, “Sparcnet: A hardware accelerator for efficient deployment of sparse convolutional networks,” ACM Journal on Emerging Technologies in Computing Systems (JETC) , vol. 13, no. 3, pp. 1–32, 2017
2017
Cited alongside, same era.
S. Yin, P. Ouyang, S. Tang, F. Tu, X. Li, S. Zheng, T. Lu, J. Gu, L. Liu, and S. Wei, “A high energy efficient reconfigurable hybrid neural network processor for deep learning applications,” IEEE Journal of Solid-State Circuits , vol. 53, no. 4, pp. 968–982, 2017
2017
Cited alongside, same era.
2019
Later among the works it cites.
M. Z. Alom, T. M. Taha, C. Yakopcic, S. Westberg, P. Sidike, M. S. Nasrin et al. , “A state-of-the-art survey on deep learning theory and architectures,” Electronics , vol. 8, no. 3, p. 292, 2019
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems , 2019, pp. 8024–8035
2019
Later among the works it cites.
K. E. Fleming, K. D. Glossop, and S. C. Steely, “Apparatus, methods, and systems with a configurable spatial accelerator,” Oct. 15 2019, uS Patent 10,445,250
2019
Later among the works it cites.
S. Dave, Y. Kim, S. Avancha, K. Lee, and A. Shrivastava, “Dmazerunner: Executing perfectly nested loops on dataflow accelerators,” ACM Transactions on Embedded Computing Systems (TECS) , vol. 18, no. 5s, pp. 1–27, 2019
2019
Later among the works it cites.
H. Kwon, P. Chatarasi, M. Pellauer, A. Parashar, V. Sarkar, and T. Krishna, “Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 754–768
2019
Later among the works it cites.
G. Srivastava, D. Kadetotad, S. Yin, V. Berisha, C. Chakrabarti, and J.-s. Seo, “Joint optimization of quantization and structured sparsity for compressed deep neural networks,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 1393–1397
2019
Later among the works it cites.
H.-J. Kang, “Accelerator-aware pruning for convolutional neural networks,” IEEE Transactions on Circuits and Systems for Video Technology , 2019
2019
Later among the works it cites.
S. Cao, L. Ma, W. Xiao, C. Zhang, Y. Liu, L. Zhang, L. Nie, and Z. Yang, “Seernet: Predicting convolutional neural network feature-map sparsity through low-bit quantization,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019
2019
Later among the works it cites.
J. Li, S. Jiang, S. Gong, J. Wu, J. Yan, G. Yan, and X. Li, “Squeezeflow: A sparse cnn accelerator exploiting concise convolution rules,” IEEE Transactions on Computers , vol. 68, no. 11, pp. 1663–1677, 2019
2019
Later among the works it cites.
J. Lee, J. Lee, D. Han, J. Lee, G. Park, and H.-J. Yoo, “7.7 lnpu: A 25.3 tflops/w sparse deep-neural-network learning processor with fine-grained mixed precision of fp8-fp16,” in 2019 IEEE International Solid-State Circuits Conference-(ISSCC) . IEEE, 2019, pp. 142–144
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
G. Georgiadis, “Accelerating convolutional neural networks via activation map compression,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 7085–7095
2019
Later among the works it cites.
U. Gupta, B. Reagen, L. Pentecost, M. Donato, T. Tambe, A. M. Rush, G.-Y. Wei, and D. Brooks, “Masr: A modular accelerator for sparse rnns,” in 2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT) . IEEE, 2019, pp. 1–14
2019
Later among the works it cites.
X. F. Xiao Dong, Lei Liu, “Acorns: A framework for accelerating deep neural networks with input sparsity,” in Proceedings of the 2019 International Conference on Parallel Architecture and Compilation (PACT) , ser. PACT ’19, 2019
2019
Later among the works it cites.
K. Hegde, H. Asghari-Moghaddam, M. Pellauer, N. Crago, A. Jaleel, E. Solomonik, J. Emer, and C. W. Fletcher, “Extensor: An accelerator for sparse tensor algebra,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
J.-F. Zhang, C.-E. Lee, C. Liu, Y. S. Shao, S. W. Keckler, and Z. Zhang, “Snap: A 1.67—21.55 tops/w sparse neural acceleration processor for unstructured sparse deep neural network inference in 16nm cmos,” in 2019 Symposium on VLSI Circuits . IEEE, 2019, pp. C306–C307
2019
Later among the works it cites.
V. Dadu, J. Weng, S. Liu, and T. Nowatzki, “Towards general purpose acceleration by exploiting common data-dependence forms,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 924–939
2019
Later among the works it cites.
J. J. Zhang, P. Raj, S. Zarar, A. Ambardekar, and S. Garg, “Compact: On-chip compression of activations for low power systolic array based cnn acceleration,” ACM Transactions on Embedded Computing Systems (TECS) , vol. 18, no. 5s, p. 47, 2019
2019
Later among the works it cites.
A. Gondimalla, N. Chesnut, M. Thottethodi, and T. Vijaykumar, “Sparten: A sparse tensor accelerator for convolutional neural networks,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture . ACM, 2019, pp. 151–165
2019
Later among the works it cites.
L. Lu, J. Xie, R. Huang, J. Zhang, W. Lin, and Y. Liang, “An efficient hardware accelerator for sparse convolutional neural networks on fpgas,” in 2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) , 2019
2019
Later among the works it cites.
B. Asgari, R. Hadidi, H. Kim, and S. Yalamanchili, “Eridanus: Efficiently running inference of dnns using systolic arrays,” IEEE Micro , vol. 39, no. 5, pp. 46–54, 2019
2019
Later among the works it cites.
H. Jang, J. Kim, J.-E. Jo, J. Lee, and J. Kim, “Mnnfast: a fast and scalable system architecture for memory-augmented neural networks,” in Proceedings of the 46th International Symposium on Computer Architecture , 2019, pp. 250–263
2019
Later among the works it cites.
B. McDanel, S. Q. Zhang, H. Kung, and X. Dong, “Full-stack optimization for accelerating cnns using powers-of-two weights with fpga validation,” in Proceedings of the ACM International Conference on Supercomputing , 2019, pp. 449–460
2019
Later among the works it cites.
2019
Later among the works it cites.
C. Hong, A. Sukumaran-Rajam, I. Nisa, K. Singh, and P. Sadayappan, “Adaptive sparse tiling for sparse matrix multiplication,” in Proceedings of the 24th Symposium on Principles and Practice of Parallel Programming . ACM, 2019, pp. 300–314
2019
Later among the works it cites.
Z. Yuan, Y. Liu, J. Yue, Y. Yang, J. Wang, X. Feng, J. Zhao, X. Li, and H. Yang, “Sticker: An energy-efficient multi-sparsity compatible accelerator for convolutional neural networks in 65-nm cmos,” IEEE Journal of Solid-State Circuits , 2019
2019
Later among the works it cites.
M. Yan, X. Hu, S. Li, A. Basak, H. Li, X. Ma, I. Akgun, Y. Feng, P. Gu, L. Deng et al. , “Alleviating irregularity in graph analytics acceleration: a hardware/software co-design approach,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 615–628
2019
Later among the works it cites.
A. Azizimazreah and L. Chen, “Shortcut mining: exploiting cross-layer shortcut reuse in dcnn accelerators,” in 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2019, pp. 94–105
2019
Later among the works it cites.
S. Sharify, A. D. Lascorz, M. Mahmoud, M. Nikolic, K. Siu, D. M. Stuart, Z. Poulos, and A. Moshovos, “Laconic deep learning inference acceleration,” in Proceedings of the 46th International Symposium on Computer Architecture . ACM, 2019, pp. 304–317
2019
Later among the works it cites.
A. Delmas Lascorz, P. Judd, D. M. Stuart, Z. Poulos, M. Mahmoud, S. Sharify, M. Nikolic, K. Siu, and A. Moshovos, “Bit-tactical: A software/hardware approach to exploiting value and bit sparsity in neural networks,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems . ACM, 2019, pp. 749–763
2019
Later among the works it cites.
A. Parashar, P. Raina, Y. S. Shao, Y.-H. Chen, V. A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A systematic approach to dnn accelerator evaluation,” in 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) . IEEE, 2019, pp. 304–315
2019
Later among the works it cites.
L. Song, J. Mao, Y. Zhuo, X. Qian, H. Li, and Y. Chen, “Hypar: Towards hybrid parallelism for deep learning accelerator array,” in 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2019, pp. 56–68
2019
Later among the works it cites.
L. R. Gonçalves, R. F. D. Moura, and L. Carro, “Aggressive energy reduction for video inference with software-only strategies,” ACM Transactions on Embedded Computing Systems (TECS) , 2019
2019
Later among the works it cites.
H. Mahdiani, A. Khadem, A. Ghanbari, M. Modarressi, F. Fattahi, and M. Daneshtalab, “ δ \delta nn: Power-efficient neural network acceleration using differential weights,” IEEE Micro , 2019
2019
Later among the works it cites.
F. Silfa, G. Dot, J.-M. Arnau, and A. Gonzàlez, “Neuron-level fuzzy memoization in rnns,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 782–793
2019
Later among the works it cites.
Y. Wang, S. Liang, H. Li, and X. Li, “A none-sparse inference accelerator that distills and reuses the computation redundancy in cnns,” in Proceedings of the 56th Annual Design Automation Conference 2019 . ACM, 2019, p. 202
2019
Later among the works it cites.
2019
Later among the works it cites.
R. Baghdadi, J. Ray, M. B. Romdhane, E. Del Sozzo, A. Akkas, Y. Zhang, P. Suriana, S. Kamil, and S. Amarasinghe, “Tiramisu: A polyhedral compiler for expressing fast and portable code,” in Proceedings of the 2019 IEEE/ACM International Symposium on Code Generation and Optimization . IEEE Press, 2019, pp. 193–205
2019
Later among the works it cites.
Y. Hu, T.-M. Li, L. Anderson, J. Ragan-Kelley, and F. Durand, “Taichi: a language for high-performance computation on spatially sparse data structures,” ACM Transactions on Graphics (TOG) , pp. 1–16, 2019
2019
Later among the works it cites.
A. Adams, K. Ma, L. Anderson, R. Baghdadi, T.-M. Li, M. Gharbi, B. Steiner, S. Johnson, K. Fatahalian, F. Durand et al. , “Learning to optimize halide with tree search and random programs,” ACM Transactions on Graphics (TOG) , vol. 38, no. 4, pp. 1–12, 2019
2019
Later among the works it cites.
C. Mendis, A. Renda, S. Amarasinghe, and M. Carbin, “Ithemal: Accurate, portable and fast basic block throughput estimation using deep neural networks,” in International Conference on Machine Learning , 2019, pp. 4505–4515
2019
Later among the works it cites.
Y. Chen, H. Lan, Z. Du, S. Liu, J. Tao, D. Han, T. Luo, Q. Guo, L. Li, Y. Xie et al. , “An instruction set architecture for machine learning,” ACM Transactions on Computer Systems (TOCS) , 2019
2019
Later among the works it cites.
S. Gopinath, N. Ghanathe, V. Seshadri, and R. Sharma, “Compiling kb-sized machine learning models to tiny iot devices,” in Proceedings of the 40th ACM SIGPLAN Conference on Programming Language Design and Implementation , 2019, pp. 79–95
2019
Later among the works it cites.
H. Yang, S. Gui, Y. Zhu, and J. Liu, “Automatic neural network compression by sparsity-quantization joint learning: A constrained optimization-based approach,” 2019
2019
Later among the works it cites.
X. Zhang, W. Jiang, Y. Shi, and J. Hu, “When neural architecture search meets hardware implementation: from hardware awareness to co-design,” in 2019 IEEE Computer Society Annual Symposium on VLSI (ISVLSI) . IEEE, 2019, pp. 25–30
2019
Later among the works it cites.
N. Srivastava, H. Rong, P. Barua, G. Feng, H. Cao, Z. Zhang et al. , “T2s-tensor: Productively generating high-performance spatial hardware for dense tensor computations,” in 2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) . IEEE, 2019, pp. 181–189
2019
Later among the works it cites.
Y.-H. Lai, Y. Chi, Y. Hu, J. Wang, C. H. Yu, Y. Zhou, J. Cong, and Z. Zhang, “Heterocl: A multi-paradigm programming infrastructure for software-defined reconfigurable computing,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 242–251
2019
Later among the works it cites.
R. Venkatesan, Y. S. Shao, M. Wang, J. Clemons, S. Dai, M. Fojtik, B. Keller, A. Klinefelter, N. Pinckney, P. Raina et al. , “Magnet: A modular accelerator generator for neural networks,” in 2019 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) , 2019
2019
Later among the works it cites.
A. Sharifian, R. Hojabr, N. Rahimi, S. Liu, A. Guha, T. Nowatzki, and A. Shriraman, “ μ \mu ir-an intermediate representation for transforming and optimizing the microarchitecture of application accelerators,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 940–953
2019
Later among the works it cites.
T. Elsken, J. H. Metzen, and F. Hutter, “Neural architecture search: A survey,” Journal of Machine Learning Research , 2019
2019
Later among the works it cites.
E. Wang, J. J. Davis, R. Zhao, H.-C. Ng, X. Niu, W. Luk, P. Y. Cheung, and G. A. Constantinides, “Deep neural network approximation for custom hardware: Where we’ve been, where we’re going,” ACM Computing Surveys (CSUR) , vol. 52, no. 2, p. 40, 2019
2019
Later among the works it cites.
A. Reuther, P. Michaleas, M. Jones, V. Gadepally, S. Samsi, and J. Kepner, “Survey and benchmarking of machine learning accelerators,” in 2019 IEEE High Performance Extreme Computing Conference (HPEC) . IEEE, 2019, pp. 1–9
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Closest in time.
N. M. Rezk, M. Purnaprajna, T. Nordström, and Z. Ul-Abdin, “Recurrent neural networks: an embedded computing perspective,” IEEE Access , vol. 8, pp. 57 967–57 996, 2020
2020
Closest in time.
2020
Closest in time.
J. Dean, “1.1 the deep learning revolution and its implications for computer architecture and chip design,” in 2020 IEEE International Solid-State Circuits Conference-(ISSCC) . IEEE, 2020, pp. 8–14
2020
Closest in time.
B. L. Deng, G. Li, S. Han, L. Shi, and Y. Xie, “Model compression and hardware acceleration for neural networks: A comprehensive survey,” Proceedings of the IEEE , vol. 108, no. 4, pp. 485–532, 2020
2020
Closest in time.
R. Krashinsky, O. Giroux, S. Jones, N. Stam, and S. Ramaswamy, “Nvidia ampere architecture in-depth,” https://devblogs.nvidia.com/nvidia-ampere-architecture-in-depth/ , 2020
2020
Closest in time.
Z.-G. Liu, P. N. Whatmough, and M. Mattina, “Systolic tensor array: An efficient structured-sparse gemm accelerator for mobile cnn inference,” IEEE Computer Architecture Letters , vol. 19, no. 1, 2020
2020
Closest in time.
E. Elsen, M. Dukhan, T. Gale, and K. Simonyan, “Fast sparse convnets,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 14 629–14 638
2020
Closest in time.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi et al. , “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, Oct. 2020, pp. 38–45
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
M. Yan, L. Deng, X. Hu, L. Liang, Y. Feng, X. Ye, Z. Zhang, D. Fan, and Y. Xie, “Hygcn: A gcn accelerator with hybrid architecture,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2020, pp. 15–29
2020
Closest in time.
T. Geng, A. Li, R. Shi, C. Wu, T. Wang, Y. Li, P. Haghi, A. Tumeo, S. Che, S. Reinhardt et al. , “Awb-gcn: A graph convolutional network accelerator with runtime workload rebalancing,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 922–936
2020
Closest in time.
H. Zeng and V. Prasanna, “Graphact: Accelerating gcn training on cpu-fpga heterogeneous platforms,” in Proceedings of the 2020 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , 2020, pp. 255–265
2020
Closest in time.
S. Liang, Y. Wang, C. Liu, L. He, L. Huawei, D. Xu, and X. Li, “Engn: A high-throughput and energy-efficient accelerator for large graph neural networks,” IEEE Transactions on Computers , 2020
2020
Closest in time.
E. Qin, A. Samajdar, H. Kwon, V. Nadella, S. Srinivasan, D. Das, B. Kaul, and T. Krishna, “Sigma: A sparse and irregular gemm accelerator with flexible interconnects for dnn training,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 58–70
2020
Closest in time.
2020
Closest in time.
X. He, S. Pal, A. Amarnath, S. Feng, D.-H. Park, A. Rovinski, H. Ye, Y. Chen, R. Dreslinski, and T. Mudge, “Sparse-tpu: Adapting systolic arrays for sparse matrices,” in Proceedings of the 34th ACM International Conference on Supercomputing , 2020, pp. 1–12
2020
Closest in time.
R. Shi, P. Dong, T. Geng, Y. Ding, X. Ma, H. K.-H. So, M. Herbordt, A. Li, and Y. Wang, “Csb-rnn: a faster-than-realtime rnn acceleration framework with compressed structured blocks,” in Proceedings of the 34th ACM International Conference on Supercomputing , 2020
2020
Closest in time.
P. Xu, X. Zhang, C. Hao, Y. Zhao, Y. Zhang, Y. Wang, C. Li, Z. Guan, D. Chen, and Y. Lin, “Autodnnchip: An automated dnn chip predictor and builder for both fpgas and asics,” in The 2020 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , 2020
2020
Closest in time.
S. Dave, A. Shrivastava, Y. Kim, S. Avancha, and K. Lee, “dmazerunner: Optimizing convolutions on dataflow accelerators,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 1544–1548
2020
Closest in time.
Intel. Understanding memory formats, intel mkl-dnn. Accessed: 2020-03-03. [Online]. Available: https://intel.github.io/mkl-dnn/understanding_memory_formats.html
2020
Closest in time.
C. Lattner, J. Pienaar, M. Amini, U. Bondhugula, R. Riddle, A. Cohen, T. Shpeisman, A. Davis, N. Vasilache, and O. Zinenko, “Mlir: A compiler infrastructure for the end of moore’s law,” 2020
2020
Closest in time.
R. Baghdadi and A. Cohen, “Scalable polyhedral compilation, syntax vs. semantics: 1–0 in the first round,” in IMPACT 2020 workshop (associated with HIPEAC 2020) , 2020, informal proceedings
2020
Closest in time.
2020
Closest in time.
Y. Wei, J. Zhou, Y. Wang, Y. Liu, Q. Liu, J. Luo, C. Wang, F. Ren, and L. Huang, “A review of algorithm & hardware design for ai-based biomedical applications.” IEEE Transactions on Biomedical Circuits and Systems , vol. 14, no. 2, pp. 145–163, 2020
2020
Closest in time.
2020
Closest in time.
B. Du, Q. Guo, Y. Zhao, T. Zhi, Y. Chen, and Z. Xu, “Self-aware neural network systems: A survey and new perspective,” Proceedings of the IEEE , 2020
2020
Closest in time.
T. Gale, M. Zaharia, C. Young, and E. Elsen, “Sparse gpu kernels for deep learning,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2020
2020
Closest in time.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” in Advances in neural information processing systems , 2016, pp. 2074–2082
2082
Closest in time.