Fetching the paper…
Reading the bibliography…
Recent works show that overparameterized networks contain small subnetworks that exhibit comparable accuracy to the full model when trained in isolation.
Optimal brain damage
Y. LeCun, J. S. Denker, and S. A. Solla · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
B. Hassibi and D. G. Stork · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Predicting parameters in deep learning
M. Denil, B. Shakibi, L. Dinh, M. A. Ranzato, and N. de Freitas · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
E. L. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus · 2014
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Y. Gong, L. Liu, M. Yang, and L. Bourdev · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
M. Jaderberg, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Speeding-up convolutional neural networks using fine-tuned cp-decomposition
V. Lebedev, Y. Ganin, M. Rakhuba, I. Oseledets, and V. Lempitsky · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
S. Han, J. Pool, J. Tran, and W. J. Dally · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Compression of deep convolutional neural networks for fast and low power mobile applications
Y.-D. Kim, E. Park, S. Yoo, T. Choi, L. Yang, and D. Shin · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Efficient and accurate approximations of nonlinear convolutional networks
X. Zhang, J. Zou, X. Ming, K. He, and J. Sun · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Structured bayesian pruning via log-normal multiplicative noise
K. Neklyudov, D. Molchanov, A. Ashukha, and D. P. Vetrov · 2017
Later among the works it cites.
Training sparse neural networks
S. Srinivas, A. Subramanya, and R. Venkatesh Babu · 2017
Later among the works it cites.
Designing energy-efficient convolutional neural networks using energy-aware pruning
T.-J. Yang, Y.-H. Chen, and V. Sze · 2017
Later among the works it cites.
The effect of network width on the performance of large-batch training
L. Chen, H. Wang, J. Zhao, D. Papailiopoulos, and P. Koutris · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Hu, R. Peng, Y.-W. Tai, and C.-K. Tang · 2016
Cited alongside, same era.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size
F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer · 2016
Cited alongside, same era.
S. Srinivas and R. V. Babu · 2016
Cited alongside, same era.
Compression-aware training of deep networks
J. M. Alvarez and M. Salzmann · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Cited alongside, same era.
Channel pruning for accelerating very deep neural networks
Y. He, X. Zhang, and J. Sun · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Cited alongside, same era.
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2018
Later among the works it cites.
Trained rank pruning for efficient deep neural networks
Y. Xu, Y. Li, S. Zhang, W. Wen, B. Wang, Y. Qi, Y. Chen, W. Lin, and H. Xiong · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Later among the works it cites.
Bag of tricks for image classification with convolutional neural networks
T. He, Z. Zhang, H. Zhang, Z. Zhang, J. Xie, and M. Li · 2019
Later among the works it cites.
The effect of network width on stochastic gradient descent and generalization: an empirical study
D. S. Park, J. Sohl-Dickstein, Q. V. Le, and S. L. Smith · 2019
Later among the works it cites.
What’s hidden in a randomly weighted neural network?
V. Ramanujan, M. Wortsman, A. Kembhavi, A. Farhadi, and M. Rastegari · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
E. Strubell, A. Ganesh, and A. McCallum · 2019
Later among the works it cites.
Scalable deep neural networks via low-rank matrix factorization
A. Yaguchi, T. Suzuki, S. Nitta, Y. Sakata, and A. Tanizawa · 2019
Later among the works it cites.
Structured binary neural networks for accurate image classification and semantic segmentation
B. Zhuang, C. Shen, M. Tan, L. Liu, and I. Reid · 2019
Later among the works it cites.
What is the state of neural network pruning?
D. Blalock, J. J. G. Ortiz, J. Frankle, and J. Guttag · 2020
Closest in time.
A low effort approach to structured cnn design using pca
I. Garg, P. Panda, and K. Roy · 2020
Closest in time.
Z. Li, E. Wallace, S. Shen, K. Lin, K. Keutzer, D. Klein, and J. E. Gonzalez · 2020
Closest in time.