Fetching the paper…
Reading the bibliography…
The parameters of a neural network are naturally organized in groups, some of which might not contribute to its overall performance.
Fonctions convexes duales et points proximaux dans un espace hilbertien
Moreau, J. J · 1962
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
Rockafellar, R. T · 1976
Earlier work this paper cites.
Regression, Prediction and Shrinkage
Copas, J. B · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Convex Analysis and Minimization Algorithms II: Advanced Theory and Bundle Methods
Hiriart-Urruty, J.-B. and Lemaréchal, C · 1993
Earlier work this paper cites.
Signal recovery by proximal forward-backward splitting
Combettes, P. L. and Wajs, V. R · 2005
Earlier work this paper cites.
Regularization and Variable Selection via the Elastic Net
Zou, H. and Hastie, T · 2005
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Yuan, M. and Lin, Y · 2006
Earlier work this paper cites.
Penalized methods for bi-level variable selection
Breheny, P. and Huang, J · 2009
Earlier work this paper cites.
Efficient Learning using Forward-Backward Splitting
Duchi, J. C. and Singer, Y · 2009
Earlier work this paper cites.
Variational Analysis
Rockafellar, R. T. and Wets, R. J.-B · 2009
Earlier work this paper cites.
Nearly Unbiased Variable Selection under Minimax Concave Penalty
Zhang, C.-H · 2010
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Convergence rates of inexact proximal-gradient methods for convex optimization
Schmidt, M., Roux, N. L., and Bach, F · 2011
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Murphy, K. P · 2012
Cited alongside, same era.
RMSprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Proximal Newton-type methods for minimizing composite functions
Lee, J. D., Sun, Y., and Saunders, M. A · 2014
Cited alongside, same era.
Proximal Algorithms
Parikh, N. and Boyd, S · 2014
Cited alongside, same era.
MADE: Masked Autoencoder for Distribution Estimation
Germain, M., Gregor, K., Murray, I., and Larochelle, H · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Later among the works it cites.
On Quasi-Newton Forward-Backward Splitting: Proximal Calculus and Convergence
Becker, S., Fadili, J., and Ochs, P · 2019
Later among the works it cites.
Reweighted Proximal Pruning for Large-Scale Language Representation
Guo, F.-M., Liu, S., Mungall, F. S., Lin, X., and Wang, Y · 2019
Later among the works it cites.
Learning Neural Causal Models from Unknown Interventions
Ke, N. R., Bilaniuk, O., Goyal, A., Bauer, S., Larochelle, H., Schölkopf, B., Mozer, M. C., Pal, C., and Bengio, Y · 2019
Later among the works it cites.
Compressed Learning of Deep Neural Networks for OpenCL-Capable Embedded Systems
Lee, S. and Lee, J · 2019
Later among the works it cites.
Toward Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kyrillidis, A., Baldassarre, L., El Halabi, M., Tran-Dinh, Q., and Cevher, V · 2015
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Simonyan, K. and Zisserman, A · 2015
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Cited alongside, same era.
Learning Structured Sparsity in Deep Neural Networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Cited alongside, same era.
Less is More: Towards Compact CNNs
Zhou, H., Alvarez, J. M., and Porikli, F · 2016
Cited alongside, same era.
Deep Roots: Improving CNN Efficiency with Hierarchical Filter Groups
Ioannou, Y., Robertson, D., Cipolla, R., and Criminisi, A · 2017
Cited alongside, same era.
Lin, S., Ji, R., Li, Y., Deng, C., and Li, X · 2019
Later among the works it cites.
Proximal Adam: Robust Adaptive Update Scheme for Constrained Optimization
Melchior, P., Joseph, R., and Moolekamp, F · 2019
Later among the works it cites.
Non-asymptotic Analysis of Stochastic Methods for Non-Smooth Non-Convex Regularized Problems
Xu, Y., Jin, R., and Yang, T · 2019
Later among the works it cites.
Half-space proximal stochastic gradient method for group-sparsity regularized problem
Chen, T., Wang, G., Ding, T., Ji, B., Yi, S., and Zhu, Z · 2020
Later among the works it cites.
Gradient-based Neural DAG Learning
Lachapelle, S., Brouillard, P., Deleu, T., and Lacoste-Julien, S · 2020
Later among the works it cites.
Proximal operator of f ( x ) = ‖ a x ‖ 2 f\left(x\right)={\left\|ax\right\|}_{2} where a a is diagonal matrix (weighted L 2 {L}_{2} norm)
Li, R · 2020
Later among the works it cites.
Movement Pruning: Adaptive Sparsity by Fine-Tuning
Sanh, V., Wolf, T., and Rush, A. M · 2020
Later among the works it cites.
ProxSGD: Training Structured Neural Networks under Regularization and Constraints
Yang, Y., Yuan, Y., Chatzimichailidis, A., van Sloun, R. J., Lei, L., and Chatzinotas, S · 2020
Later among the works it cites.
A General Family of Stochastic Proximal Gradient Methods for Deep Learning
Yun, J., Lozano, A. C., and Yang, E · 2020
Later among the works it cites.