Fetching the paper…
Reading the bibliography…
Transformer is a transformative framework that models sequential data and has achieved remarkable performance on a wide range of tasks, but with high computational and energy cost.
Algorithm 778: L-bfgs-b: Fortran subroutines for large-scale bound-constrained optimization
C. Zhu, R. H. Byrd, P. Lu, and J. Nocedal · 1997
Earlier work this paper cites.
Similarity search in high dimensions via hashing
A. Gionis, P. Indyk, and R. Motwani · 1999
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
M. S. Charikar · 2002
Earlier work this paper cites.
Locality-sensitive hashing scheme based on p-stable distributions
M. Datar, N. Immorlica, P. Indyk, and V. S. Mirrokni · 2004
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Earlier work this paper cites.
Spectral hashing
Y. Weiss, A. Torralba, and R. Fergus · 2008
Earlier work this paper cites.
Kernelized locality-sensitive hashing
B. Kulis and K. Grauman · 2011
Earlier work this paper cites.
Hashing with graphs
W. Liu, J. Wang, S. Kumar, and S.-F. Chang · 2011
Earlier work this paper cites.
Minimal loss hashing for compact binary codes
M. Norouzi and D. J. Fleet · 2011
Earlier work this paper cites.
Iterative quantization: A procrustean approach to learning binary codes for large-scale image retrieval
Y. Gong, S. Lazebnik, A. Gordo, and F. Perronnin · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Supervised hashing with kernels
W. Liu, J. Wang, R. Ji, Y.-G. Jiang, and S.-F. Chang · 2012
Earlier work this paper cites.
Semi-supervised hashing for large-scale search
J. Wang, S. Kumar, and S.-F. Chang · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons
Y. Bengio · 2013
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
M. Horowitz · 2014
Earlier work this paper cites.
Fast supervised hashing with decision trees for high-dimensional data
G. Lin, C. Shen, Q. Shi, A. Van den Hengel, and D. Suter · 2014
Earlier work this paper cites.
Discrete graph hashing
W. Liu, C. Mu, S. Kumar, and S.-F. Chang · 2014
Earlier work this paper cites.
Practical and optimal lsh for angular distance
A. Andoni, P. Indyk, T. Laarhoven, I. Razenshteyn, and L. Schmidt · 2015
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
M. Courbariaux, Y. Bengio, and J.-P. David · 2015
Earlier work this paper cites.
Eie: Efficient inference engine on compressed deep neural network
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally · 2016
Earlier work this paper cites.
Binarized neural networks
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou · 2016
Earlier work this paper cites.
Fast training of triplet-based deep binary embedding networks
B. Zhuang, G. Lin, C. Shen, and I. Reid · 2016
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2017
Cited alongside, same era.
Towards accurate binary convolutional neural network
X. Lin, C. Zhao, and W. Pan · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Loss-aware weight quantization of deep networks
L. Hou and J. T. Kwok · 2018
Cited alongside, same era.
Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm
Z. Liu, B. Wu, W. Luo, X. Yang, W. Liu, and K.-T. Cheng · 2018
Cited alongside, same era.
Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy
Q-bert: Hessian based ultra low precision quantization of bert
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer · 2020
Later among the works it cites.
Fast transformers with clustered attention
A. Vyas, A. Katharopoulos, and F. Fleuret · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
S. Wang, B. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Later among the works it cites.
Deep multi-view enhancement hashing for image retrieval
C. Yan, B. Gong, Y. Wei, and Y. Gao · 2020
Later among the works it cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
A. H. Zadeh, I. Edo, O. M. Awad, and A. Moshovos · 2020
Later among the works it cites.
Binarybert: Pushing the limit of bert quantization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Mishra and D. Marr · 2018
Cited alongside, same era.
Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network
H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, V. Chandra, and H. Esmaeilzadeh · 2018
Cited alongside, same era.
Proxquant: Quantized neural networks via proximal operators
Y. Bai, Y.-X. Wang, and E. Liberty · 2019
Cited alongside, same era.
Xnor-net++: Improved binary neural networks
A. Bulat and G. Tzimiropoulos · 2019
Cited alongside, same era.
Universal transformers
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Regularizing activation distribution for training binarized deep networks
R. Ding, T.-W. Chin, Z. Liu, and D. Marculescu · 2019
Cited alongside, same era.
H. Bai, W. Zhang, L. Hou, L. Shang, J. Jin, X. Jiang, Q. Liu, M. Lyu, and I. King · 2021
Later among the works it cites.
Rethinking attention with performers
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller · 2021
Later among the works it cites.
Twins: Revisiting the design of spatial attention in vision transformers
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Soft: Softmax-free transformer with linear complexity
J. Lu, J. Yao, J. Zhang, X. Zhu, H. Xu, W. Gao, C. Xu, T. Xiang, and L. Zhang · 2021
Later among the works it cites.
Random feature attention
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. Smith, and L. Kong · 2021
Later among the works it cites.
Combiner: Full attention transformer with sparse computation cost
H. Ren, H. Dai, Z. Dai, M. Yang, J. Leskovec, D. Schuurmans, and B. Dai · 2021
Later among the works it cites.
Adder attention for vision transformer
H. Shu, J. Wang, H. Chen, L. Li, Y. Yang, and Y. Wang · 2021
Later among the works it cites.
Long range arena: A benchmark for efficient transformers
Y. Tay, M. Dehghani, S. Abnar, Y. Shen, D. Bahri, P. Pham, J. Rao, L. Yang, S. Ruder, and D. Metzler · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
Nyströmformer: A nystöm-based algorithm for approximating self-attention
Y. Xiong, Z. Zeng, R. Chakraborty, M. Tan, G. Fung, Y. Li, and V. Singh · 2021
Later among the works it cites.
Long-short transformer: Efficient transformers for language and vision
C. Zhu, W. Ping, C. Xiao, M. Shoeybi, T. Goldstein, A. Anandkumar, and B. Catanzaro · 2021
Later among the works it cites.
BiBERT: Accurate fully binarized BERT
H. Qin, Y. Ding, M. Zhang, Q. YAN, A. Liu, Q. Dang, Z. Liu, and X. Liu · 2022
Closest in time.
cosformer: Rethinking softmax in attention
Z. Qin, W. Sun, H. Deng, D. Li, Y. Wei, B. Lv, J. Yan, L. Kong, and Y. Zhong · 2022
Closest in time.
Sparse attention with learning to hash
Z. Sun, Y. Yang, and S. Yoo · 2022
Closest in time.
Pvtv2: Improved baselines with pyramid vision transformer
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2022
Closest in time.