Fetching the paper…
Reading the bibliography…
Popular deep learning frameworks require users to fine-tune their memory usage so that the training data of a deep neural network (DNN) fits within the GPU physical memory.
A. Robinson and C. Cherry, “Results of a Prototype Television Bandwidth Compression Scheme,” Proceedings of the IEEE
1967
Earlier work this paper cites.
S. Hanson and L. Pratt, “Comparing Biases for Minimal Network Construction with Back-propagation,” in Proceedings of the International Conference on Neural Information Processing Systems (NIPS)
1989
Earlier work this paper cites.
Y. LeCun, S. Denker, and S. Solla, “Optimal Brain Damage,” in Proceedings of the International Conference on Neural Information Processing Systems (NIPS)
1990
Earlier work this paper cites.
B. Hassibi and D. Stork, “Second Order Derivatives for Network Pruning: Optimal Brain Surgeon,” in Proceedings of the International Conference on Neural Information Processing Systems (NIPS)
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long Short Term Memory,” Neural Computation
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-Based Learning Applied to Document Recognition,” Proceedings of the IEEE
1998
Earlier work this paper cites.
P. Wilson, S. Kaplan, and Y. Smaragdakis, “The Case for Compressed Cache in Virtual Memory Systems,” in Proceedings of USENIX
1999
Earlier work this paper cites.
Y. Zhang, J. Yang, and R. Gupta, “Frequent Value Locality and Value-centric Data Cache Design,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS)
2000
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Framewise Phoneme Classification With Bidirectional LSTM and Other Neural Network Architectures,” Neural Networks
2005
Earlier work this paper cites.
H. Wong, M. M. Papadopoulou, M. Sadooghi-Alvandi, and A. Moshovos, “Demystifying GPU Microarchitecture Through Microbenchmarking,” in Proceedings of the International Symposium on Performance Analysis of Systems Software (ISPASS)
2010
Earlier work this paper cites.
2011
Earlier work this paper cites.
V. Vanhoucke, A. Senior, and M. Mao, “Improving the Speed of Neural Networks on CPUs,” in Deep Learning and Unsupervised Feature Learning Workshop
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Proceedings of the International Conference on Neural Information Processing Systems (NIPS)
2012
Earlier work this paper cites.
A. Krizhevsky, “cuda-convnet.” https://code.google.com/p/cuda-convnet/ , 2012
2012
Earlier work this paper cites.
V. Sathish, M. Schulte, and N. Kim, “Lossless and Lossy Memory I/O Link Compression for Improving Performance of GPGPU Workloads,” in Proceedings of the International Conference on Parallel Architectures and Compilation Techniques (PACT)
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. Lin, Q. Chen, and S. Yan, “Network in Network.” https://arxiv.org/abs/1312.4400 , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer, “cuDNN: Efficient Primitives for Deep Learning,” in Proceedings of the International Conference on Neural Information Processing Systems (NIPS)
2014
Earlier work this paper cites.
B. Pichai, L. Hsu, and A. Bhattacharjee, “Architectural Support for Address Translation on GPUs: Designing Memory Management Units for CPU/GPUs with Unified Address Spaces,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS)
2014
Earlier work this paper cites.
J. Power, M. Hill, and D. Wood, “Supporting x86-64 Address Translation for 100s of GPU Lanes,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Abdelfattah, A. Hagiescu, and D. Singh, “Gzip on a Chip: High Performance Lossless Data Compression on FPGAs Using OpenCL,” in Proceedings of the International Workshop on OpenCL
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Research
2014
Cited alongside, same era.
2014
Cited alongside, same era.
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “DianNao: A Small-footprint High-throughput Accelerator for Ubiquitous Machine-learning,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS)
2014
Cited alongside, same era.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “DaDianNao: A Machine-Learning Supercomputer,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2014
Cited alongside, same era.
M. Rhu, N. Gimelshein, J. Clemons, A. Zulfiqar, and S. W. Keckler, “vDNN: Virtualized Deep Neural Networks for Scalable, Memory-Efficient Neural Network Design,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2016
Later among the works it cites.
2016
Later among the works it cites.
T. Zheng, D. Nellans, A. Zulfiqar, M. Stephenson, and S. W. Keckler, “Toward High-Performance Paged-Memory for GPUs,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2016
Later among the works it cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-Neuron-Free Deep Convolutional Neural Network Computing,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang, “MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems,” in Proceedings of the Workshop on Machine Learning Systems
2015
Cited alongside, same era.
2015
Cited alongside, same era.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going Deeper with Convolutions,” in Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Y. Sun, X. Wang, and X. Tang, “Deeply Learned Face Representations Are Sparse, Selective, and Robust,” in Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
ImageNet. http://image-net.org , 2016
2016
Later among the works it cites.
G. Diamos, S. Sengupta, B. Catanzaro, M. Chrzanowski, A. Coates, E. Elsen, J. Engel, A. Hannun, and S. Satheesh, “Persistent RNNs: Stashing Recurrent Weights On-Chip,” in Proceedings of the International Conference on Machine Learning (ICML)
2016
Later among the works it cites.
2016
Later among the works it cites.
G. Pekhimenko, E. Bolotin, N. Vijaykumar, O. Mutlu, T. C. Mowry, and S. W. Keckler, “A Case for Toggle-Aware Compression for GPU Systems,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2016
Later among the works it cites.
NCSU, “FreePDK Process Design Kit.” http://www.eda.ncsu.edu/wiki/FreePDK , 2016
2016
Later among the works it cites.
HP Labs, “CACTI: An Integrated Cache and Memory Access Time, Cycle Time, Area, Leakage, and Dynamic Power Model.” http://www.hpl.hp.com/research/cacti/ , 2016
2016
Later among the works it cites.
“GPGPU-Sim.” http://www.gpgpu-sim.org/ , 2016
2016
Later among the works it cites.
“GPU Ocelot: A Dynamic Compilation Framework for GPU Computing.” http://gpuocelot.gatech.edu/ , 2016
2016
Later among the works it cites.
S. Han, H. Mao, and W. Dally, “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” in Proceedings of the International Conference on Learning Representations (ICLR)
2016
Later among the works it cites.
Y. Chen, T. Krishna, J. Emer, and V. Sze, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” in Proceedings of the International Solid State Circuits Conference (ISSCC)
2016
Later among the works it cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. Horowitz, and W. Dally, “EIE: Efficient Inference Engine on Compressed Deep Neural Network,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Later among the works it cites.
Y. Chen, J. Emer, and V. Sze, “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Later among the works it cites.
R. LiKamWa, Y. Hou, M. Polansky, Y. Gao, and L. Zhong, “RedEye: Analog ConvNet Image Sensor Architecture for Continuous Mobile Vision,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Later among the works it cites.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. Lee, J. Miguel, H. Lobato, G. Wei, and D. Brooks, “Minerva: Enabling Low-Power, High-Accuracy Deep Neural Network Accelerators,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Later among the works it cites.
P. Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “A Novel Processing-in-memory Architecture for Neural Network Computation in ReRAM-based Main Memory,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Later among the works it cites.
A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V. Srikumar, “ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Later among the works it cites.
NVIDIA, “NVIDIA NVLINK High-Speed Interconnect.” http://www.nvidia.com/object/nvlink.html , 2016
2016
Later among the works it cites.
IBM, “IBM Power Systems.” http://www-03.ibm.com/systems/power/ , 2016
2016
Later among the works it cites.
NVIDIA, “The NVIDIA DGX-1 Deep Learning System.” http://www.nvidia.com/object/deep-learning-system.html , 2016
2016
Later among the works it cites.