Fetching the paper…
Reading the bibliography…
There has been an explosion of interest in designing high-performance Transformers.
P. Indyk and R. Motwani, “Approximate nearest neighbors: Towards removing the curse of dimensionality,” in Proceedings of the Thirtieth Annual ACM Symposium on the Theory of Computing, Dallas, Texas, USA, May 23-26, 1998 . ACM, 1998, pp. 604–613
1998
Earlier work this paper cites.
H. Jegou, M. Douze, and C. Schmid, “Product quantization for nearest neighbor search,” PAMI , no. 1, pp. 117–128, 2010
2010
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. W. Senior, P. A. Tucker, K. Yang, and A. Y. Ng, “Large scale distributed deep networks,” in NIPS , 2012, pp. 1232–1240
2012
Earlier work this paper cites.
M. Courbariaux, Y. Bengio, and J. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” in NIPS , 2015, pp. 3123–3131
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” IJCV , pp. 211–252, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv: Learning , 2016
2016
Earlier work this paper cites.
J. Ba, J. Kiros, and G. E. Hinton, “Layer normalization,” ArXiv , vol. abs/1607.06450, 2016
2016
Earlier work this paper cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “Xnor-net: Imagenet classification using binary convolutional neural networks,” in ECCV , 2016, pp. 525–542
2016
Earlier work this paper cites.
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “QSGD: communication-efficient SGD via gradient quantization and encoding,” in NIPS , 2017, pp. 1709–1720
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “Terngrad: Ternary gradients to reduce communication in distributed deep learning,” in NIPS , 2017, pp. 1509–1519
2017
Earlier work this paper cites.
P. Micikevicius, S. Narang, J. Alben, G. F. Diamos, E. Elsen, D. García, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in ICLR , 2018
2018
Earlier work this paper cites.
J. Choi, Z. Wang, S. Venkataramani, P. I.-J. Chuang, V. Srinivasan, and K. Gopalakrishnan, “PACT: Parameterized clipping activation for quantized neural networks,” 2018. [Online]. Available: https://openreview.net/forum?id=By5ugjyCb
2018
Earlier work this paper cites.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. G. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in CVPR , 2018, pp. 2704–2713
2018
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in ICML , 2019, pp. 6105–6114
2019
Cited alongside, same era.
H. Mostafa and X. Wang, “Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization,” in ICML , 2019, pp. 4646–4655
2019
Cited alongside, same era.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. X. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, and Z. Chen, “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” in NIPS , 2019, pp. 103–112
2019
Cited alongside, same era.
A. Chakrabarti and B. Moseley, “Backprop with approximate activations for memory-efficient network training,” in NIPS , 2019, pp. 2426–2435
2019
Cited alongside, same era.
S. Jung, C. Son, S. Lee, J. Son, J. Han, Y. Kwak, S. J. Hwang, and C. Choi, “Learning to quantize deep networks by optimizing quantization intervals with task loss,” in CVPR , 2019, pp. 4350–4359
S. K. Esser, J. L. McKinstry, D. Bablani, R. Appuswamy, and D. S. Modha, “Learned step size quantization,” in ICLR , 2020
2020
Later among the works it cites.
P. Wang, Q. Chen, X. He, and J. Cheng, “Towards accurate post-training network quantization via bit-split and stitching,” in ICML , 2020, pp. 9847–9856
2020
Later among the works it cites.
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Q-BERT: hessian based ultra low precision quantization of BERT,” in AAAI , 2020, pp. 8815–8821
2020
Later among the works it cites.
W. Zhang, L. Hou, Y. Yin, L. Shang, X. Chen, X. Jiang, and Q. Liu, “Ternarybert: Distillation-aware ultra-low bit BERT,” in EMNLP , 2020, pp. 509–521
2020
Later among the works it cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in ICCV , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
M. Nagel, M. van Baalen, T. Blankevoort, and M. Welling, “Data-free quantization through weight equalization and bias correction,” in ICCV , 2019
2019
Cited alongside, same era.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat, “Q8BERT: quantized 8bit BERT,” in EMC2@NeurIPS , 2019, pp. 36–39
2019
Cited alongside, same era.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in NIPS , 2019, pp. 8024–8035
2019
Cited alongside, same era.
K. Wang, Z. Liu, Y. Lin, J. Lin, and S. Han, “HAQ: hardware-aware automated quantization with mixed precision,” in CVPR , 2019, pp. 8612–8620
2019
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Cited alongside, same era.
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ADE20K dataset,” IJCV , pp. 302–321, 2019
2019
Cited alongside, same era.
A. Kirillov, R. B. Girshick, K. He, and P. Dollár, “Panoptic feature pyramid networks,” in CVPR , 2019, pp. 6399–6408
2019
Cited alongside, same era.
2021
Closest in time.
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. S. Torr, and L. Zhang, “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in CVPR , 2021, pp. 6881–6890
2021
Closest in time.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021
2021
Closest in time.
J. Chen, L. Zheng, Z. Yao, D. Wang, I. Stoica, M. W. Mahoney, and J. Gonzalez, “Actnn: Reducing training memory footprint via 2-bit activation compressed training,” in ICML , 2021, pp. 1803–1813
2021
Closest in time.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” ICLR , 2021
2021
Closest in time.
Z. Pan, B. Zhuang, J. Liu, H. He, and J. Cai, “Scalable visual transformers with hierarchical pooling,” in ICCV , 2021
2021
Closest in time.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” in ICCV , 2021
2021
Closest in time.
B. Chen, P. Li, C. Li, B. Li, L. Bai, C. Lin, M. Sun, J. Yan, and W. Ouyang, “Glit: Neural architecture search for global and local image transformer,” in ICCV , 2021
2021
Closest in time.
M. Chen, H. Peng, J. Fu, and H. Ling, “Autoformer: Searching transformers for visual recognition,” in ICCV , 2021
2021
Closest in time.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in ICML , 2021
2021
Closest in time.
H. Bai, W. Zhang, L. Hou, L. Shang, J. Jin, X. Jiang, Q. Liu, M. R. Lyu, and I. King, “Binarybert: Pushing the limit of BERT quantization,” in ACL/IJCNLP , 2021, pp. 4334–4348
2021
Closest in time.
M. Croci, M. Fasi, N. Higham, T. Mary, and M. Mikaitis, “Stochastic rounding: Implementation, error analysis, and applications,” 2021
2021
Closest in time.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” in NeurIPS 2021 , 2021
2021
Closest in time.