Fetching the paper…
Reading the bibliography…
Sparsification of neural networks is one of the effective complexity reduction methods to improve efficiency and generalizability.
Annals of Mathematical Statistics 22
Robbins, H., Monro, S.: A stochastic approximation method · 1951
Earlier work this paper cites.
USSR Computational Mathematics and Mathematical Physics 4
Polyak, B.: Some methods of speeding up the convergence of iteration methods · 1964
Earlier work this paper cites.
Nature 323
Rumelhart, D., Hinton, G., Williams, R.: Learning representations by back-propagating errors · 1986
Earlier work this paper cites.
SIAM Journal on Applied Mathematics 61
Nikolova, M.: Local strong homogeneity of a regularized estimator · 2000
Earlier work this paper cites.
In Advances in neural information processing systems pp. 342–350 (2009)
Cho, Y., Saul, L.K.: Kernel methods for deep learning · 2009
Earlier work this paper cites.
Journal of Machine Learning Research 12
Duchi, J., Hazan, E., Singer, Y.: Adaptive subgradient methods for online learning and stochastic optimization · 2011
Earlier work this paper cites.
IEEE Signal Processing Magazine 29
Hinton, G., Deng, L., Yu, D., Dahl, G.E., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T.N., Kingsbury, B.: Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups · 2012
Earlier work this paper cites.
In Advances in neural information processing systems pp. 1097–1105 (2012)
Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks · 2012
Earlier work this paper cites.
ArXiv preprint 1412.6980
Kingma, D., Ba, J.: Adam: A method for stochastic optimization. (2014) · 2014
Earlier work this paper cites.
Optimization Methods and Software 30
Lu, Z., Zhang, Y.: Penalty decomposition methods for rank minimization · 2014
Cited alongside, same era.
ArXiv preprint 1510.00149
Han, S., Mao, H., Dally, W.J.: Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding (2015) · 2015
Cited alongside, same era.
ArXiv preprint 1612.08083
Dauphin, Y.N., Fan, A., Auli, M., Grangier, D.: Language modeling with gated convolutional networks. (2016) · 2016
Cited alongside, same era.
ArXiv preprint 1609.01037
Shamir, O.: Distribution-specific hardness of learning neural networks (2016) · 2016
Cited alongside, same era.
In: International Conference on Machine Learning, pp. 2722–2731 (2016)
Taylor, G., Burmeister, R., Xu, Z., Singh, B., Patel, A., Goldstein, T.: Training neural networks without gradients: A scalable admm approach · 2016
Cited alongside, same era.
ArXiv preprint 1611.03530
Zhang, C., Bengio, S., Hardt, M., Recht, B., Vinyals, O.: Understanding deep learning requires rethinking generalization (2016) · 2016
ArXiv preprint 1703.00560
Tian, Y.: An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis (2017) · 2017
Later among the works it cites.
Tech. rep., Technical report (2017)
Tieleman, T., Hinton, G.: Divide the gradient by a running average of its recent magnitude. coursera: Neural networks for machine learning · 2017
Later among the works it cites.
ICLR (2017)
Ullrich, K., Meeds, E., Welling, M.: Soft weight-sharing for neural network compression · 2017
Later among the works it cites.
Communications in Mathematical Sciences 15
Zhang, S., Xin, J.: Minimization of transformed l 1 l_{1} penalty: Closed form representation and iterative thresholding algorithms · 2017
Later among the works it cites.
In: International Conference on Machine Learning (ICML) (2018)
Du, S., Lee, J., Tian, Y., Poczos, B., Singh, A.: Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima · 2018
Closest in time.
ArXiv preprint 1712.01312v2
Louizos, C., Welling, M., Kingma, D.: Learning sparse neural networks through ℓ 0 \ell_{0} regularization (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ArXiv preprint 1702.07966
Brutzkus, A., Globerson, A.: Globally optimal gradient descent for a convnet with gaussian inputs (2017) · 2017
Cited alongside, same era.
ArXiv 1709.06129
Du, S., Lee, J., Tian, Y.: When is a convolutional filter easy to learn? (2017) · 2017
Cited alongside, same era.
ArXiv preprint 1701.05369
Molchanov, D., Ashukha, A., Vetrov, D.: Variational dropout sparsifies deep neural networks (2017) · 2017
Cited alongside, same era.
Closest in time.
In: International Conference on Learning Representations (2018)
Reddi, S., Kale, S., Kumar, S.: On the convergence of adam and beyond · 2018
Closest in time.
Journal of Scientific Computing, online (2018)
Wang, Y., Zeng, J., Yin, W.: Global Convergence of ADMM in Nonconvex Nonsmooth Optimization · 2018
Closest in time.
arXiv preprint 1804.03294 (2018)
Zhang, T., Ye, S., Zhang, K., Tang, J., Wen, W., Fardad, M., Wang, Y.: A systematic dnn weight pruning framework using alternating direction method of multipliers · 2018
Closest in time.