Fetching the paper…
Reading the bibliography…
Recent research has focused on weight sparsity in deep neural network training to reduce FLOPs, aiming for improved efficiency (test accuracy w.r.t training FLOPs).
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
Robust principal component analysis?
Candès, E. J., Li, X., Ma, Y., and Wright, J · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Convolutional neural networks with low-rank regularization
Tai, C., Xiao, T., Zhang, Y., Wang, X., and Weinan, E · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
DeVries, T. and Taylor, G. W · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Pyramid scene parsing network
Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Earlier work this paper cites.
Encoder-decoder with atrous separable convolution for semantic image segmentation
Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H · 2018
Earlier work this paper cites.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D., Mocanu, E., Stone, P., Nguyen, P., Gibescu, M., and Liotta, A · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Earlier work this paper cites.
MMDetection: Open mmlab detection toolbox and benchmark
Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Xu, J., Zhang, Z., Cheng, D., Zhu, C., Cheng, T., Zhao, Q., Li, B., Lu, X., Zhu, R., Wu, Y., Dai, J., Wang, J., Shi, J., Ouyang, W., Loy, C. C., and Lin, D · 2019
Earlier work this paper cites.
Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks
Ding, X., Guo, Y., Ding, G., and Han, J · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
Bag of tricks for image classification with convolutional neural networks
He, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., and Li, M · 2019
Cited alongside, same era.
Resnet v1.5 for pytorch
Nvidia · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
WINOGRANDE: an adversarial winograd schema challenge at scale, 2019
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Cited alongside, same era.
EfficientNet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Cited alongside, same era.
Why are big data matrices approximately low rank?
Udell, M. and Townsend, A · 2019
Cited alongside, same era.
Thinking outside the die: Architecting the ml accelerator of the future
Lie, S · 2021
Later among the works it cites.
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer
Mehta, S. and Rastegari, M · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2021
Later among the works it cites.
Deepsparse, 2021
NeuralMagic · 2021
Later among the works it cites.
Scalable vision transformers with hierarchical pooling
Pan, Z., Zhuang, B., Liu, J., He, H., and Cai, J · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H., Lan, J., Liu, R., and Yosinski, J · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
The lottery ticket hypothesis for pre-trained bert networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Cited alongside, same era.
MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark
Contributors, M · 2020
Cited alongside, same era.
Progressive skeletonization: Trimming more fat from a network at initialization
de Jorge, P., Sanyal, A., Behl, H. S., Torr, P. H., Rogez, G., and Dokania, P. K · 2020
Cited alongside, same era.
Later among the works it cites.
Bottleneck transformers for visual recognition
Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A · 2021
Later among the works it cites.
Search spaces for neural model training
Stosic, D. and Stosic, D · 2021
Later among the works it cites.
Doping: A technique for efficient compression of lstm models using sparse structured additive matrices
Thakker, U., Whatmough, P. N., Liu, Z., Mattina, M., and Beu, J · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
Wu, H., Xiao, B., Codella, N., Liu, M., Dai, X., Yuan, L., and Zhang, L · 2021
Later among the works it cites.
Mest: Accurate and fast memory-economic sparse training framework on the edge
Yuan, G., Ma, X., Niu, W., Li, Z., Kong, Z., Liu, N., Gong, Y., Zhan, Z., He, C., Jin, Q., et al · 2021
Later among the works it cites.
Do-conv: Depthwise over-parameterized convolutional layer
Cao, J., Li, Y., Sun, M., Chen, Y., Lischinski, D., Cohen-Or, D., Chen, B., and Tu, C · 2022
Later among the works it cites.
Sparsity winning twice: Better robust generalization from more efficient training
Chen, T., Zhang, Z., pengjun wang, Balachandra, S., Ma, H., Wang, Z., and Wang, Z · 2022
Later among the works it cites.
Monarch: Expressive structured matrices for efficient and accurate training
Dao, T., Chen, B., Sohoni, N. S., Desai, A., Poli, M., Grogan, J., Liu, A., Rao, A., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Vision models are more robust and fair when pretrained on uncurated images without supervision
Goyal, P., Duval, Q., Seessel, I., Caron, M., Singh, M., Misra, I., Sagun, L., Joulin, A., and Bojanowski, P · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., Sifre, L., and et al · 2022
Later among the works it cites.
Exposing and exploiting fine-grained block structures for fast and accurate sparse training
Jiang, P., Hu, L., and Song, S · 2022
Later among the works it cites.
Truthfulqa: Measuring how models mimic human falsehoods, 2022
Lin, S., Hilton, J., and Evans, O · 2022
Later among the works it cites.
Effective model sparsification by scheduled grow-and-prune methods
Ma, X., Qin, M., Sun, F., Hou, Z., Yuan, K., Xu, Y., Wang, Y., Chen, Y.-K., Jin, R., and Xie, Y · 2022
Later among the works it cites.
Spartan: Differentiable Sparsity via Regularized Transportation
Tai, K. S., Tian, T., and Lim, S.-N · 2022
Later among the works it cites.
An improved one millisecond mobile backbone
Vasu, P. K. A., Gabriel, J., Zhu, J., Tuzel, O., and Ranjan, A · 2022
Later among the works it cites.
Getting vit in shape: Scaling laws for compute-optimal model design
Alabdulmohsin, I., Zhai, X., Kolesnikov, A., and Beyer, L · 2023
Closest in time.
Open llm leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T · 2023
Closest in time.
Train a model with weight sparsity
Cerebras · 2023
Closest in time.
Dynamic sparse training via balancing the exploration-exploitation trade-off
Huang, S., Lei, B., Xu, D., Peng, H., Sun, Y., Xie, M., and Ding, C · 2023
Closest in time.
Cerebras architecture deep dive: First look inside the hardware/software co-design for deep learning
Lie, S · 2023
Closest in time.
Nvidia performance documentation
Nvidia · 2023
Closest in time.
SPDF: Sparse pre-training and dense fine-tuning for large language models
Thangarasa, V., Gupta, A., Marshall, W., Li, T., Leong, K., DeCoste, D., Lie, S., and Saxena, S · 2023
Closest in time.