Fetching the paper…
Reading the bibliography…
We provide a new efficient version of the backpropagation algorithm, specialized to the case where the weights of the neural network being trained are sparse.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
A simple and effective method for removal of hidden units and weights
Hagiwara, M · 1994
Earlier work this paper cites.
Learning generative visual models from few training examples: an incremental Bayesian approach tested on 101 object categories
Li, F.-F., Fergus, R., and Perona, P · 2004
Earlier work this paper cites.
The Caltech 256
Griffin, G., Holub, A. D., and Perona, P · 2006
Earlier work this paper cites.
A visual vocabulary for flower classification
Nilsback, M.-E. and Zisserman, A · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J., Hays, J., Ehinger, K., Oliva, A., and Torralba, A · 2010
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V · 2012
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S., Rahtu, E., Kannala, J., Blaschko, M., and Vedaldi, A · 2013
Earlier work this paper cites.
Birdsnap: Large-scale fine-grained visual categorization of birds
Berg, T., Liu, J., Lee, S. W., Alexander, M. L., Jacobs, D. W., and Belhumeur, P. N · 2014
Earlier work this paper cites.
Food-101 – mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Van Gool, L · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
pybind11 — seamless operability between c++11 and python, 2016
Jakob, W., Rhinelander, J., and Moldovan, D · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
A topological insight into restricted boltzmann machines
Mocanu, D. C., Mocanu, E., Nguyen, P. H., Gibescu, M., and Liotta, A · 2016
Earlier work this paper cites.
Faster cnns with direct sparse convolutions and guided pruning
Park, J., Li, S., Wen, W., Tang, P. T. P., Li, H., Chen, Y., and Dubey, P · 2016
Cited alongside, same era.
Gpu kernels for block-sparse weights
Gray, S., Radford, A., and Kingma, D. P · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Cited alongside, same era.
Deep rewiring: Training very sparse deep networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Sparsert: Accelerating unstructured sparsity on gpus for deep learning inference
Wang, Z · 2020
Later among the works it cites.
Dithered backprop: A sparse and quantized backpropagation algorithm for more efficient deep neural network training
Wiedemann, S., Mehari, T., Kepp, K., and Samek, W · 2020
Later among the works it cites.
Procrustes: a dataflow and accelerator for sparse deep neural network training
Yang, D., Ghasemazar, A., Ren, X., Golub, M., Lemieux, G., and Lis, M · 2020
Later among the works it cites.
Memorized sparse backpropagation
Zhang, Z., Yang, P., Ren, X., Su, Q., and Sun, X · 2020
Later among the works it cites.
The lottery tickets hypothesis for supervised and self-supervised pre-training in computer vision models
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Carbin, M., and Wang, Z · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sparse networks from scratch: Faster training without losing performance
Dettmers, T. and Zettlemoyer, L · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Do better imagenet models transfer better?
Kornblith, S., Shlens, J., and Le, Q. V · 2019
Cited alongside, same era.
Full deep neural network training on a pruned weight budget
Lis, M., Golub, M., and Lemieux, G · 2019
Cited alongside, same era.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Mostafa, H. and Wang, X · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Cited alongside, same era.
Fast sparse convnets
Elsen, E., Dukhan, M., Gale, T., and Simonyan, K · 2020
Cited alongside, same era.
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Later among the works it cites.
Accelerated sparse neural training: A provable and efficient method to find N:M transposable masks
Hubara, I., Chmiel, B., Island, M., Banner, R., Naor, S., and Soudry, D · 2021
Later among the works it cites.
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Lhoest, Q., Villanova del Moral, A., Jernite, Y., Thakur, A., von Platen, P., Patil, S., Chaumond, J., Drame, M., Plu, J., Tunstall, L., Davison, J., Šaško, M., Chhablani, G., Malik, B., Brandeis, S., Le Scao, T., Sanh, V., Xu, C., Patry, N., McMillan-Major, A., Schmid, P., Gugger, S., Delangue, C., Matussière, T., Debut, L., Bekman, S., Cistac, P., Goehringer, T., Mustar, V., Lagunas, F., Rush, A., and Wolf, T · 2021
Later among the works it cites.
Accelerating sparse deep neural networks
Mishra, A., Latorre, J. A., Pool, J., Stosic, D., Stosic, D., Venkatesh, G., Yu, C., and Micikevicius, P · 2021
Later among the works it cites.
Accelerating Inference with Sparsity Using the NVIDIA Ampere Architecture and NVIDIA TensorRT, 2021
NVIDIA · 2021
Later among the works it cites.
AC/DC: Alternating compressed/decompressed training of deep neural networks
Peste, A., Iofinova, E., Vladu, A., and Alistarh, D · 2021
Later among the works it cites.
Powerpropagation: A sparsity inducing weight reparameterisation
Schwarz, J., Jayakumar, S., Pascanu, R., Latham, P., and Teh, Y · 2021
Later among the works it cites.
Prune once for all: Sparse pre-trained language models
Zafrir, O., Larey, A., Boudoukh, G., Shen, H., and Wasserblat, M · 2021
Later among the works it cites.
How well do sparse ImageNet models transfer?
Iofinova, E., Peste, A., Kurtz, M., and Alistarh, D · 2022
Later among the works it cites.
Sten: An interface for efficient sparsity in pytorch
Ivanov, A., Dryden, N., and Hoefler, T · 2022
Later among the works it cites.
Exposing and exploiting fine-grained block structures for fast and accurate sparse training
Jiang, P., Hu, L., and Song, S · 2022
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Later among the works it cites.
DeepSparse, 2022
NeuralMagic · 2022
Later among the works it cites.
DeepSparse, 2022
SparseZoo, N · 2022
Later among the works it cites.