Fetching the paper…
Reading the bibliography…
We examine the performance profile of Convolutional Neural Network training on the current generation of NVIDIA Graphics Processing Units.
An algorithm for the machine calculation of complex fourier series
Cooley, James W. and Tukey, John W · 1965
Earlier work this paper cites.
A linear filtering approach to the computation of discrete Fourier transform
Bluestein, Leo I · 1970
Earlier work this paper cites.
Supernode partitioning
Irigoin, F. and Triolet, R · 1988
Earlier work this paper cites.
A bridging model for parallel computation
Valiant, Leslie G · 1990
Earlier work this paper cites.
Understanding Digital Signal Processing
Lyons, Richard G · 1996
Earlier work this paper cites.
A family of high-performance matrix multiplication algorithms
Gunnels, John A., Henry, Greg M., and van de Geijn, Robert A · 2001
Earlier work this paper cites.
High Performance Convolutional Neural Networks for Document Processing
Chellapilla, Kumar, Puri, Sidd, and Simard, Patrice · 2006
Earlier work this paper cites.
Fast fourier transforms, 2008
Burrus, C. Sidney · 2008
Earlier work this paper cites.
Parallel computing experiences with cuda
Garland, Michael, Le Grand, Scott, Nickolls, John, Anderson, Joshua, Hardwick, Jim, Morton, Scott, Phillips, Everett, Zhang, Yao, and Volkov, Vasily · 2008
Cited alongside, same era.
Optimizing matrix transpose in cuda
Ruetsch, Greg and Micikevicius, Paulius · 2009
Cited alongside, same era.
Theano: a CPU and GPU math expression compiler
Bergstra, James, Breuleux, Olivier, Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Desjardins, Guillaume, Turian, Joseph, Warde-Farley, David, and Bengio, Yoshua · 2010
Cited alongside, same era.
Speeding up nek5000 with autotuning and specialization
Shin, Jaewook, Hall, Mary W., Chame, Jacqueline, Chen, Chun, Fischer, Paul F., and Hovland, Paul D · 2010
Cited alongside, same era.
Better performance at lower occupancy
Volkov, V · 2010
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
cudnn: Efficient primitives for deep learning
Chetlur, Sharan, Woolley, Cliff, Vandermersch, Philippe, Cohen, Jonathan, Tran, John, Catanzaro, Bryan, and Shelhamer, Evan · 2014
Closest in time.
Course on cuda programming on nvidia gpus, lecture 3, 2014
Giles, Mike · 2014
Closest in time.
Caffe: Convolutional architecture for fast feature embedding
Jia, Yangqing, Shelhamer, Evan, Donahue, Jeff, Karayev, Sergey, Long, Jonathan, Girshick, Ross, Guadarrama, Sergio, and Darrell, Trevor · 2014
Closest in time.
cuda-convnet2, 2014
Krizhevsky, Alex · 2014
Closest in time.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Sermanet, Pierre, Eigen, David, Zhang, Xiang, Mathieu, Michael, Fergus, Rob, and LeCun, Yann · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast training of convolutional networks through ffts
Mathieu, Michaël, Henaff, Mikael, and LeCun, Yann · 2013
Cited alongside, same era.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Ragan-Kelley, Jonathan, Barnes, Connelly, Adams, Andrew, Paris, Sylvain, Durand, Frédo, and Amarasinghe, Saman P · 2013
Cited alongside, same era.
Torch7: A matlab-like environment for machine learning
Collobert, R., Kavukcuoglu, K., and Farabet, C
Cited in the paper.
Natural language processing (almost) from scratch
Collobert, Ronan, Weston, Jason, Bottou, Léon, Karlen, Michael, Kavukcuoglu, Koray, and Kuksa, Pavel
Cited in the paper.
Efficient training of convolutional deep belief networks in the frequency domain for application to high-resolution 2d and 3d images
Brosch, Tom and Tam, Roger C · 2015
Closest in time.
maxdnn: An efficient convolution kernel for deep learning with maxwell gpus
Lavin, Andrew · 2015
Closest in time.