Fetching the paper…
Reading the bibliography…
In this paper, we introduce a new perspective on training deep neural networks capable of state-of-the-art performance without the need for the expensive over-parameterization by proposing the concept of In-Time Over-Parameterization (ITOP) in sparse training.
Pruning versus clipping in neural networks
Janowsky, S. A · 1989
Earlier work this paper cites.
Using relevance to reduce network size automatically
Mozer, M. C. and Smolensky, P · 1989
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D. G · 1993
Earlier work this paper cites.
Pruning neural networks at initialization: Why are we missing the mark?
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2009
Earlier work this paper cites.
Gradient flow in sparse neural networks and how lottery tickets win
Evci, U., Ioannou, Y. A., Keskin, C., and Dauphin, Y · 2010
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Dynamic network surgery for efficient dnns
Guo, Y., Yao, A., and Chen, Y · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
A topological insight into restricted boltzmann machines
Mocanu, D. C., Mocanu, E., Nguyen, P. H., Gibescu, M., and Liotta, A · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D. and Carmon, Y · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Earlier work this paper cites.
Runtime neural pruning
Lin, J., Rao, Y., Lu, J., and Zhou, J · 2017
Earlier work this paper cites.
Learning sparse neural networks through l _ 0 l\_0 regularization
Louizos, C., Welling, M., and Kingma, D. P · 2017
Earlier work this paper cites.
Network computations in artificial intelligence
Mocanu, D. C · 2017
Earlier work this paper cites.
Variational dropout sparsifies deep neural networks
Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Earlier work this paper cites.
Exploring sparsity in recurrent neural networks
Narang, S., Elsen, E., Diamos, G., and Sengupta, S · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., Ying, C., and Le, Q. V · 2017
Cited alongside, same era.
Training sparse neural networks
Srinivas, S., Subramanya, A., and Venkatesh Babu, R · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Cited alongside, same era.
Deep rewiring: Training very sparse deep networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Autoprune: Automatic network pruning by regularizing auxiliary parameters
Xiao, X., Wang, Z., and Rajasekaran, S · 2019
Later among the works it cites.
Drawing early-bird tickets: Towards more efficient training of deep networks
You, H., Li, C., Xu, P., Fu, Y., Wang, Y., Chen, X., Baraniuk, R. G., Wang, Z., and Lin, Y · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Zou, D. and Gu, Q · 2019
Later among the works it cites.
Atashgahi, Z., Sokar, G., van der Lee, T., Mocanu, E., Mocanu, D. C., Veldhuis, R., and Pechenizkiy, M · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
Nest: A neural network synthesis tool based on a grow-and-prune paradigm
Dai, X., Yin, H., and Jha, N. K · 2019
Cited alongside, same era.
Sparse networks from scratch: Faster training without losing performance
Dettmers, T. and Zettlemoyer, L · 2019
Cited alongside, same era.
Later among the works it cites.
Progressive skeletonization: Trimming more fat from a network at initialization
de Jorge, P., Sanyal, A., Behl, H. S., Torr, P. H., Rogez, G., and Dokania, P. K · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Top-kast: Top-k always sparse training
Jayakumar, S., Pascanu, R., Rae, J., Osindero, S., and Elsen, E · 2020
Later among the works it cites.
Soft threshold weight reparameterization for learnable sparsity
Kusupati, A., Ramanujan, V., Somani, R., Wortsman, M., Jain, P., Kakade, S., and Farhadi, A · 2020
Later among the works it cites.
A signal propagation perspective for pruning neural networks at initialization
Lee, N., Ajanthan, T., Gould, S., and Torr, P. H. S · 2020
Later among the works it cites.
Sparse weight activation training
Raihan, M. A. and Aamodt, T. M · 2020
Later among the works it cites.
An investigation of why overparameterization exacerbates spurious correlations
Sagawa, S., Raghunathan, A., Koh, P. W., and Liang, P · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H., Kunin, D., Yamins, D. L., and Ganguli, S · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Wang, C., Zhang, G., and Grosse, R · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2020
Later among the works it cites.
Progressive skeletonization: Trimming more fat from a network at initialization
de Jorge, P., Sanyal, A., Behl, H., Torr, P., Rogez, G., and Dokania, P. K · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Closest in time.
Liu, S., Mocanu, D. C., Pei, Y., and Pechenizkiy, M · 2021
Closest in time.
Leveraging sparse linear layers for debuggable deep networks
Wong, E., Santurkar, S., and Madry, A · 2021
Closest in time.
Can subnetwork structure be the key to out-of-distribution generalization?, 2021
Zhang, D., Ahuja, K., Xu, Y., Wang, Y., and Courville, A · 2021
Closest in time.
Learning n: M fine-grained structured sparse neural networks from scratch
Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H · 2021
Closest in time.