Fetching the paper…
Reading the bibliography…
Convolutional Neural Networks (CNNs) are the state of the art solution for many computer vision problems, and many researchers have explored optimized implementations.
J. J. Dongarra, J. Du Croz, S. Hammarling, and I. S. Duff, “A set of level 3 basic linear algebra subprograms,” ACM Transactions on Mathematical Software (TOMS) , vol. 16, no. 1, pp. 1–17, 1990
1990
Earlier work this paper cites.
S. Browne, J. Dongarra, N. Garner, G. Ho, and P. Mucci, “A portable programming interface for performance evaluation on modern processors,” International Journal of High Performance Computing Applications , vol. 14, no. 3, pp. 189–204, 2000
2000
Earlier work this paper cites.
J. A. Gunnels, G. M. Henry, and R. A. van de Geijn, “A family of high-performance matrix multiplication algorithms,” in Proceedings of the International Conference on Computational Sciences-Part I , ser. ICCS ’01. London, UK, UK: Springer-Verlag, 2001, pp. 51–60
2001
Earlier work this paper cites.
R. C. Whaley, A. Petitet, and J. J. Dongarra, “Automated empirical optimizations of software and the atlas project,” Parallel Computing , vol. 27, no. 1, pp. 3–35, 2001
2001
Earlier work this paper cites.
M. Frigo and V. Strumpen, “Cache oblivious stencil computations,” in Proceedings of the 19th annual international conference on Supercomputing . ACM, 2005, pp. 361–366
2005
Earlier work this paper cites.
G. Hinton, S. Osindero, and Y.-W. Teh, “A fast learning algorithm for deep belief nets,” Neural computation , vol. 18, no. 7, pp. 1527–1554, 2006
2006
Earlier work this paper cites.
M. T. Inc., “TN-41-01: Calculating Memory System Power for DDR3,” http://www.micron.com/support/power-calc , 2007
2007
Earlier work this paper cites.
N. Muralimanohar, R. Balasubramonian, and N. Jouppi, “Optimizing NUCA Organizations and Wiring Alternatives for Large Caches with CACTI 6.0,” in Proceedings of the 40th Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO 40. Washington, DC, USA: IEEE Computer Society, 2007, pp. 3–14
2007
Earlier work this paper cites.
U. Bondhugula, A. Hartono, J. Ramanujam, and P. Sadayappan, “A practical automatic polyhedral parallelizer and locality optimizer,” in Proceedings of the 29th ACM SIGPLAN Conference on Programming Language Design and Implementation , ser. PLDI ’08. New York, NY, USA: ACM, 2008, pp. 101–113
2008
Earlier work this paper cites.
K. Goto and R. A. van de Geijn, “Anatomy of high-performance matrix multiplication,” ACM Trans. Math. Softw. , vol. 34, no. 3, pp. 12:1–12:25, May 2008
2008
Earlier work this paper cites.
K. Jarrett, K. Kavukcuoglu, M. Ranzato, and Y. LeCun, “What is the best multi-stage architecture for object recognition?” in 12th Int’l Conf. on Computer Vision . IEEE, 2009, pp. 2146–2153
2009
Earlier work this paper cites.
J. Bergstra, F. Bastien, J. Turian, R. Pascanu, O. Delalleau, O. Breuleux, P. Lamblin, G. Desjardins, D. Erhan, Y. Bengio et al. , “Deep learning on GPUs with Theano,” in The Learning Workshop-Research Abstract-Oral preferred (Feb. 18, 2010) , 2010
2010
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Convolutional deep belief networks on CIFAR-10,” Unpublished manuscript , 2010
2010
Earlier work this paper cites.
D. Strigl, K. Kofler, and S. Podlipnig, “Performance and scalability of GPU-based convolutional neural networks,” in Parallel, Distributed and Network-Based Processing (PDP), 2010 18th Euromicro International Conference on , Feb 2010, pp. 317–324
2010
Earlier work this paper cites.
R. Strzodka, M. Shaheen, D. Pajak, and H.-P. Seidel, “Cache oblivious parallelograms in iterative stencil computations,” in Proceedings of the 24th ACM International Conference on Supercomputing . ACM, 2010, pp. 49–59
2010
Cited alongside, same era.
S. C. Turaga, J. F. Murray, V. Jain, F. Roth, M. Helmstaedter, K. Briggman, W. Denk, and H. S. Seung, “Convolutional networks can learn to generate affinity graphs for image segmentation,” Neural Computation , vol. 22, no. 2, pp. 511–538, 2010
2010
Cited alongside, same era.
C. Farabet, B. Martini, B. Corda, P. Akselrod, E. Culurciello, and Y. LeCun, “NeuFlow: a runtime reconfigurable dataflow processor for vision,” in IEEE Computer Society Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW) . IEEE, 2011, pp. 109–116
2011
Cited alongside, same era.
P. Sermanet and Y. LeCun, “Traffic sign recognition with multi-scale convolutional networks,” in Neural Networks (IJCNN), The 2011 International Joint Conference on . IEEE, 2011, pp. 2809–2813
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun et al. , “DaDianNao: a machine-learning supercomputer,” in 47th Annual IEEE/ACM Int’l Symp. on Microarchitecture (MICRO) . IEEE, 2014, pp. 609–622
2014
Later among the works it cites.
A. Dundar, J. Jin, V. Gokhale, B. Martini, and E. Culurciello, “Memory access optimized routing scheme for deep networks on a mobile coprocessor,” Algorithms , vol. 12, p. 15, 2014
2014
Later among the works it cites.
V. Gokhale, J. Jin, A. Dundar, B. Martini, and E. Culurciello, “A 240 G-ops/s mobile coprocessor for deep neural networks,” in IEEE Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW) . IEEE, 2014, pp. 696–701
2014
Later among the works it cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in Proceedings of the ACM International Conference on Multimedia . ACM, 2014, pp. 675–678
2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25 , F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds. Curran Associates, Inc., 2012, pp. 1097–1105
2012
Cited alongside, same era.
P.-H. Pham, D. Jelaca, C. Farabet, B. Martini, Y. LeCun, and E. Culurciello, “NeuFlow: dataflow vision processing system-on-a-chip,” in 55th Int’l Midwest Symp. on Circuits and Systems (MWSCAS) . IEEE, 2012, pp. 1044–1047
2012
Cited alongside, same era.
J. Ragan-Kelley, A. Adams, S. Paris, M. Levoy, S. P. Amarasinghe, and F. Durand, “Decoupling algorithms from schedules for easy optimization of image processing pipelines.” ACM Trans. Graph. , vol. 31, no. 4, p. 32, 2012
2012
Cited alongside, same era.
M. Peemen, A. A. Setio, B. Mesman, and H. Corporaal, “Memory-centric accelerator design for convolutional neural networks,” in 31st Int’l Conf. on Computer Design (ICCD) . IEEE, 2013, pp. 13–19
2013
Cited alongside, same era.
L.-N. Pouchet, P. Zhang, P. Sadayappan, and J. Cong, “Polyhedral-based data reuse optimization for configurable computing,” in Proceedings of the ACM/SIGDA international symposium on Field programmable gate arrays . ACM, 2013, pp. 29–38
2013
Cited alongside, same era.
J. Ragan-Kelley, C. Barnes, A. Adams, S. Paris, F. Durand, and S. Amarasinghe, “Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines,” ACM SIGPLAN Notices , vol. 48, no. 6, pp. 519–530, 2013
2013
Cited alongside, same era.
D. Sanchez and C. Kozyrakis, “Zsim: Fast and accurate microarchitectural simulation of thousand-core systems,” in Proceedings of the 40th Annual International Symposium on Computer Architecture , ser. ISCA ’13. New York, NY, USA: ACM, 2013, pp. 475–486
2013
Cited alongside, same era.
U. Bondhugula, V. Bandishti, A. Cohen, G. Potron, and N. Vasilache, “Tiling and optimizing time-iterated computations on periodic domains,” in Proceedings of the 23rd international conference on Parallel architectures and compilation . ACM, 2014, pp. 39–50
2014
Cited alongside, same era.
Later among the works it cites.
J. Jin, V. Gokhale, A. Dundar, B. Krishnamurthy, B. Martini, and E. Culurciello, “An efficient implementation of deep convolutional neural networks on a mobile coprocessor,” in IEEE 57th Int’l Midwest Symp. on Circuits and Systems (MWSCAS) . IEEE, 2014, pp. 133–136
2014
Later among the works it cites.
NVIDIA, “cuDNN,” https://developer.nvidia.com/cuDNN , 2014
2014
Later among the works it cites.
2014
Later among the works it cites.
E. Wang, Q. Zhang, B. Shen, G. Zhang, X. Lu, Q. Wu, and Y. Wang, “Intel math kernel library,” in High-Performance Computing on the Intel Xeon Phi . Springer, 2014, pp. 167–188
2014
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
D. Liu, T. Chen, S. Liu, J. Zhou, S. Zhou, O. Teman, X. Feng, X. Zhou, and Y. Chen, “PuDianNao: a polyvalent machine learning accelerator,” in Proc. 20th Int’l Conf. on Architectural Support for Programming Languages and Operating Systems . ACM, 2015, pp. 369–381
2015
Later among the works it cites.
2015
Later among the works it cites.
U. Bondhugula, A. Acharya, and A. Cohen, “The pluto+ algorithm: A practical approach for parallelization and locality optimization of affine loop nests,” ACM Trans. On Programming Languages and Systems (TOPLAS) , 2016
2016
Closest in time.