Fetching the paper…
Reading the bibliography…
The most widely used machine learning frameworks require users to carefully tune their memory usage so that the deep neural network (DNN) fits into the DRAM capacity of a GPU.
S. Hanson and L. Pratt, “Comparing Biases for Minimal Network Construction with Back-propagation,” in Proceedings of the Advances in Neural Information Processing Systems
1989
Earlier work this paper cites.
Y. LeCun, S. Denker, and S. Solla, “Optimal Brain Damage,” in Proceedings of the Advances in Neural Information Processing Systems
1990
Earlier work this paper cites.
B. Hassibi and D. Stork, “Second Order Derivatives for Network Pruning: Optimal Brain Surgeon,” in Proceedings of the Advances in Neural Information Processing Systems
1993
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-Based Learning Applied to Document Recognition,” in Proceedings of the IEEE
1998
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Framewise Phoneme Classification With Bidirectional LSTM and Other Neural Network Architectures,” in Neural Networks
2005
Earlier work this paper cites.
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa, “Natural Language Processing (Almost) From Scratch,” in arxiv.org
2011
Earlier work this paper cites.
V. Vanhoucke, A. Senior, and M. Mao, “Improving the Speed of Neural Networks on CPUs,” in Proceedings of Deep Learning and Unsupervised Feature Learning
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Proceedings of the Advances in Neural Information Processing Systems
2012
Earlier work this paper cites.
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks,” in arxiv.org
2013
Earlier work this paper cites.
OpenMP Architecture Review Board, “OpenMP Application Program Interface (version 4.0),” 2013
2013
Earlier work this paper cites.
A. Krizhevsky, “One Weird Trick For Parallelizing Convolutional Neural Networks,” in arxiv.org
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going Deeper with Convolutions,” in arxiv.org
2014
Earlier work this paper cites.
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning,” in Proceedings of International Conference on Architectural Support for Programming Languages and Operating Systems
2014
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “DaDianNao: A Machine-Learning Supercomputer,” in Proceedings of ACM/IEEE International Symposium on Microarchitecture
2014
Earlier work this paper cites.
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer, “cuDNN: Efficient Primitives for Deep Learning,” in Proceedings of the Advances in Neural Information Processing Systems
2014
Earlier work this paper cites.
Y. Gong, L. Liu, M. Yang, and L. Bourdev, “Compressing Deep Convolutional Networks Using Vector Quantization,” in arxiv.org
2014
Earlier work this paper cites.
B. Pichai, L. Hsu, and A. Bhattacharjee, “Architectural Support for Address Translation on GPUs: Designing Memory Management Units for CPU/GPUs with Unified Address Spaces,” in Proceedings of ACM International Conference on Architectural Support for Programming Languages and Operating Systems
2014
Earlier work this paper cites.
J. Power, M. Hill, and D. Wood, “Supporting x86-64 Address Translation for 100s of GPU Lanes,” in Proceedings of IEEE International Symposium on High-Performance Computer Architecture
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in Proceedings of the International Conference on Learning Representations
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in arxiv.org
2015
Cited alongside, same era.
S. Chintala, “https://github.com/torch/nn/pull/235,” 2015
2015
Cited alongside, same era.
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang, “MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems,” in Proceedings of the 2015 Workshop on Machine Learning Systems
2015
Cited alongside, same era.
NVIDIA, “GeForce GTX Titan X (Maxwell),” 2015
2015
Cited alongside, same era.
Y. Chen, T. Krishna, J. Emer, and V. Sze, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” in IEEE International Conference on Solid-State Circuits
2016
Closest in time.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. Horowitz, and W. Dally, “EIE: Efficient Inference Engine on Compressed Deep Neural Network,” in Proceedings of ACM/IEEE International Symposium on Computer Architecture
2016
Closest in time.
Y. Chen, J. Emer, and V. Sze, “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” in Proceedings of ACM/IEEE International Symposium on Computer Architecture
2016
Closest in time.
R. LiKamWa, Y. Hou, M. Polansky, Y. Gao, and L. Zhong, “RedEye: Analog ConvNet Image Sensor Architecture for Continuous Mobile Vision,” in Proceedings of ACM/IEEE International Symposium on Computer Architecture
2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Han, J. Pool, J. Tran, and W. Dally, “Learning Both Weights and Connections for Efficient Neural Networks,” in Proceedings of the Advances in Neural Information Processing Systems
2015
Cited alongside, same era.
Z. Du, R. Fasthuber, T. Chen, P. Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam, “ShiDianNao: Shifting Vision Processing Closer to the Sensor,” in Proceedings of ACM/IEEE International Symposium on Computer Architecture
2015
Cited alongside, same era.
Tensorflow, “https://www.tensorflow.org,” 2016
2016
Cited alongside, same era.
Torch, “http://torch.ch,” 2016
2016
Cited alongside, same era.
Theano, “http://deeplearning.net/tutorial,” 2016
2016
Cited alongside, same era.
Caffe, “http://caffe.berkeleyvision.org,” 2016
2016
Cited alongside, same era.
NVIDIA, “cuDNN: GPU Accelerated Deep Learning,” 2016
2016
Cited alongside, same era.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. Lee, J. Miguel, H. Lobato, G. Wei, and D. Brooks, “Minerva: Enabling Low-Power, High-Accuracy Deep Neural Network Accelerators,” in Proceedings of ACM/IEEE International Symposium on Computer Architecture
2016
Closest in time.
P. Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “A Novel Processing-in-memory Architecture for Neural Network Computation in ReRAM-based Main Memory,” in Proceedings of ACM/IEEE International Symposium on Computer Architecture
2016
Closest in time.
A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V. Srikumar, “ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” in Proceedings of ACM/IEEE International Symposium on Computer Architecture
2016
Closest in time.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-Neuron-Free Deep Convolutional Neural Network Computing,” in Proceedings of ACM/IEEE International Symposium on Computer Architecture
2016
Closest in time.
T. Zheng, D. Nellans, A. Zulfiqar, M. Stephenson, and S. W. Keckler, “Toward High-Performance Paged-Memory for GPUs,” in Proceedings of IEEE International Symposium on High-Performance Computer Architecture
2016
Closest in time.
NVIDIA, “NVIDIA NVLINK High-Speed Interconnect,” 2016
2016
Closest in time.
NVIDIA, “NVIDIA CUDA Programming Guide,” 2016
2016
Closest in time.
NVIDIA, “https://github.com/NVIDIA/cnmem,” 2016
2016
Closest in time.
S. Gross and M. Wilber, “Training and Investigating Residual Nets,” 2016
2016
Closest in time.
S. Chintala, “https://github.com/soumith/convnet-benchmarks,” 2016
2016
Closest in time.
NVIDIA, “CUDA Toolkit 7.5 Documentation: Profiler,” 2016
2016
Closest in time.
S. Han, H. Mao, and W. Dally, “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” in Proceedings of the International Conference on Learning Representations
2016
Closest in time.
P. Judd, J. Albericio, T. Hetherington, T. Aamodt, N. E. Jerger, R. Urtasun, and A. Moshovos, “Reduced-Precision Strategies for Bounded Memory in Deep Neural Nets,” in arxiv.org
2016
Closest in time.