Fetching the paper…
Reading the bibliography…
Memory is increasingly often the bottleneck when training neural network models.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Algorithm 799: Revolve: An implementation of checkpointing for the reverse or adjoint mode of computational differentiation
Andreas Griewank and Andrea Walther · 2000
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
cuSPARSE, 2011
NVIDIA Corporation · 2011
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Tara N. Sainath, B. Kingsbury, Vikas Sindhwani, E. Arisoy, and B. Ramabhadran · 2013
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Earlier work this paper cites.
Compressing neural networks with the hashing trick
Wenlin Chen, James T. Wilson, Stephen Tyree, Killian Q. Weinberger, and Yixin Chen · 2015
Earlier work this paper cites.
An exploration of parameter redundancy in deep networks with circulant projections
Yu Cheng, Felix X. Yu, Rogerio S. Feris, Sanjiv Kumar, Alok Choudhary, and Shih-Fu Chang · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
GPU-based deep learning inference: A performance and power analysis
NVIDIA Corporation · 2015
Earlier work this paper cites.
Structured transforms for small-footprint deep learning
Vikas Sindhwani, Tara N. Sainath, and Sanjiv Kumar · 2015
Earlier work this paper cites.
Deep fried convnets
Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alex Smola, Le Song, and Ziyu Wang · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Earlier work this paper cites.
EIE: efficient inference engine on compressed deep neural network
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and < 0.5 <0.5 MB model size
Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
Qianli Liao, Kenji Kawaguchi, and Tomaso Poggio · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2016
Earlier work this paper cites.
In-place activated batchnorm for memory-optimized training of DNNs
Samuel Rota Bulò, Lorenzo Porzi, and Peter Kontschieder · 2017
Earlier work this paper cites.
CirCNN: Accelerating and compressing deep neural networks using block-circulant weight matrices
Caiwen Ding, Siyu Liao, Yanzhi Wang, Zhe Li, Ning Liu, Youwei Zhuo, Chao Wang, Xuehai Qian, Yu Bai, Geng Yuan, Xiaolong Ma, Yipeng Zhang, Jian Tang, Qinru Qiu, Xue Lin, and Bo Yuan · 2017
Earlier work this paper cites.
NVIDIA Tesla V100 GPU architecture, 2017
Luke Durant, Olivier Giroux, Mark Harris, and Nick Stam · 2017
Earlier work this paper cites.
The reversible residual network: Backpropagation without storing activations
Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse · 2017
Earlier work this paper cites.
Learning spatio-temporal features with 3D residual networks for action recognition
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
MobileNets: Efficient convolutional neural networks for mobile vision applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Sergey Ioffe · 2017
Cited alongside, same era.
CNN-based segmentation of medical imaging data
Barış Kayalıbay, Grady Jensen, and Patrick van der Smagt · 2017
Cited alongside, same era.
Sharan Narang, Gregory Diamos, Erich Elsen, Paulius Micikevicius, Jonah Alben, David Garcia, Boris Ginsburg, Michael Houston, Ganesh Venkatesh, and Hao Wu · 2017
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal C. Mocanu, Elena Mocanu, Peter Stone, Phuong H. Nguyen, Madeleine Gibescu, and Antonio Liotta · 2018
Later among the works it cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Later among the works it cites.
Language models are unsupervised multitask learners, 2018
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2018
Later among the works it cites.
Compressing deep neural networks with probabilistic data structures
Brandon Reagen, Udit Gupta, Robert Adolf, Michael M. Mitzenmacher, Alexander M. Rush, Gu-Yeon Wei, and David Brooks · 2018
Later among the works it cites.
Compressing DMA engine: Leveraging activation sparsity for training deep neural networks
Minsoo Rhu, Mike O’Connor, Niladrish Chatterjee, Jeff Pool, Youngeun Kwon, and Stephen W. Keckler · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
SCNN: An accelerator for compressed-sparse convolutional neural networks
Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, and William J. Dally · 2017
Cited alongside, same era.
Faster CNNs with direct sparse convolutions and guided pruning
Jongsoo Park, Sheng Li, Wei Wen, Ping Tak Peter Tang, Hai Li, Yiran Chen, and Pradeep Dubey · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Plasticine: A reconfigurable architecture for parallel patterns
Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, Stefan Hadjis, Ardavan Pedram, Christos Kozyrakis, and Kunle Olukotun · 2017
Cited alongside, same era.
Learning spatio-temporal representation with pseudo-3D residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Cited alongside, same era.
Inception-v4, Inception-ResNet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Cited alongside, same era.
MobileNetV2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Later among the works it cites.
Learning compressed transforms with low displacement rank
Anna Thomas, Albert Gu, Tri Dao, Atri Rudra, and Christopher Ré · 2018
Later among the works it cites.
HAQ: hardware-aware automated quantization
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han · 2018
Later among the works it cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Speeding up ImageNet training on supercomputers
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Later among the works it cites.
Memory-efficient adaptive optimization for large-scale learning
Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer · 2019
Closest in time.
moDNN: Memory optimal deep neural network training on graphics processing units
Xiaoming Chen, Danny Ziyi Chen, Yinhe Han, and Xiaobo Sharon Hu · 2019
Closest in time.
Accurate and efficient 2-bit quantized neural networks
Jungwook Choi, Swagath Venkataramani, Vijayalakshmi Srinivasan, Kailash Gopalakrishnan, Zhuo Wang, and Pierce Chuang · 2019
Closest in time.
Learning fast algorithms for linear transforms using butterfly factorizations
Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra, and Christopher Ré · 2019
Closest in time.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and MIchael Carbin · 2019
Closest in time.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Closest in time.
Full deep neural network training on a pruned weight budget
Maximilian Golub, Guy Lemieux, and Mieszko Lis · 2019
Closest in time.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Yonglong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, and Zhifeng Chen · 2019
Closest in time.
Sambhav R. Jain, Albert Gural, Michael Wu, and Chris Dick · 2019
Closest in time.
Dynamic sparse graph for efficient deep learning
Liu Liu, Lei Deng, Xing Hu, Maohua Zhu, Guoqi Li, Yufei Ding, and Yuan Xie · 2019
Closest in time.
Differentiable learning-to-normalize via switchable normalization
Ping Luo, Jiamin Ren, Zhanglin Peng, Ruimao Zhang, and Jingyu Li · 2019
Closest in time.
Mini-batch serialization: CNN training with inter-layer data reuse
Sangkug Lym, Armand Behroozi, Wei Wen, Ge Li, Yongkee Kwon, and Mattan Erez · 2019
Closest in time.
Traditional and heavy-tailed self regularization in neural network models
Charles H. Martin and Michael W. Mahoney · 2019
Closest in time.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Hesham Mostafa and Xin Wang · 2019
Closest in time.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N. Dauphin, and Michael Auli · 2019
Closest in time.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Closest in time.
Buddy compression: Enabling larger memory for deep learning and HPC workloads on GPUs
Esha Choukse, Michael B. Sullivan, Mike O’Connor, Mattan Erez, Jeff Pool, David Nellans, and Stephen W. Keckler · 2020
Closest in time.
Sparse evolutionary deep learning with over one million artificial neurons on commodity hardware
Shiwei Liu, Decebal Constantin Mocanu, Amarsagar Reddy Ramapuram Matavalam, Yulong Pei, and Mykola Pechenizkiy · 2021
Closest in time.