Fetching the paper…
Reading the bibliography…
Transformers are the backbone of powerful foundation models for many Vision and Natural Language Processing tasks.
Optimal brain damage
Lecun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Hassibi, B., Stork, D. G., and Wolff, G. J · 1993
Earlier work this paper cites.
Data compression and harmonic analysis
Donoho, D., Vetterli, M., DeVore, R., and Daubechies, I · 1998
Earlier work this paper cites.
Quantized overcomplete expansions in R n {R}^{n} : analysis, synthesis, and algorithms
Goyal, V., Vetterli, M., and Thao, N · 1998
Earlier work this paper cites.
Quantized frame expansions with erasures
Goyal, V. K., Kovačević, J., and Kelner, J. A · 2000
Earlier work this paper cites.
Equal-norm tight frames with erasures
Casazza, P. G. and Kovačević, J · 2003
Earlier work this paper cites.
Grassmannian frames with applications to coding and communication
Strohmer, T. and Heath Jr, R. W · 2003
Earlier work this paper cites.
Analysis of noise reduction in redundant expansions under distributed processing requirements
Rozell, C. and Johnson, D · 2005
Earlier work this paper cites.
Fusion frames and distributed processing
Casazza, P. G., Kutyniok, G., and Li, S · 2008
Earlier work this paper cites.
Beyond bandlimited sampling: Nonlinearities, smoothness and sparsity
Eldar, Y. C. and Michaeli, T · 2008
Earlier work this paper cites.
Robust dimension reduction, fusion frames, and grassmannian packings
Kutyniok, G., Pezeshki, A., Calderbank, R., and Liu, T · 2008
Earlier work this paper cites.
Compressed sensing for fusion frames
Boufounos, P., Kutyniok, G., and Rauhut, H · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Constructing tight fusion frames
Casazza, P. G., Fickus, M., Mixon, D. G., Wang, Y., and Zhou, Z · 2010
Earlier work this paper cites.
Finite frames, theory and applications, 2012
Casazza, P. G. and Kutyniok, G · 2012
Earlier work this paper cites.
Exploiting sparseness in deep neural networks for large vocabulary speech recognition
Yu, D., Seide, F., Li, G., and Deng, L · 2012
Earlier work this paper cites.
Fastfood-approximating kernel expansions in loglinear time
Le, Q., Sarlós, T., Smola, A., et al · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Shallowing deep networks: Layer-wise pruning based on feature representations
Chen, S. and Zhao, Q · 2018
Earlier work this paper cites.
An introduction to frames and riesz bases, 2018
Christensen, O · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Cited alongside, same era.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Banner, R., Nahshan, Y., and Soudry, D · 2019
Cited alongside, same era.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Cited alongside, same era.
Data-free quantization through weight equalization and bias correction
Nagel, M., Baalen, M. v., Blankevoort, T., and Welling, M · 2019
Cited alongside, same era.
An introduction to finite tight frames, 2019
Waldron, S. F. D · 2019
Cited alongside, same era.
Towards accurate post-training quantization for vision transformer
Ding, Y., Qin, H., Yan, Q., Chai, Z., Liu, J., Wei, X., and Liu, X · 2022
Later among the works it cites.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Frantar, E. and Alistarh, D · 2022
Later among the works it cites.
A survey of quantization methods for efficient neural network inference
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K · 2022
Later among the works it cites.
Deit iii: Revenge of the vit
Touvron, H., Cord, M., and Jégou, H · 2022
Later among the works it cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Yazdani Aminabadi, R., Zhang, M., et al · 2022
Later among the works it cites.
Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pytorch image models
Wightman, R · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Cited alongside, same era.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., et al · 2020
Cited alongside, same era.
Up or down? Adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
O (n) connections are expressive enough: Universal approximability of sparse transformers
Yun, C., Chang, Y.-W., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Cited alongside, same era.
Yuan, Z., Xue, C., Chen, Y., Wu, Q., and Sun, G · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Later among the works it cites.
QuIP: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D · 2023
Later among the works it cites.
fast-hadamard-transform, 2023
Dao, T · 2023
Later among the works it cites.
Harmonic grassmannian codes
Fickus, M., Iverson, J. W., Jasper, J., and Mixon, D. G · 2023
Later among the works it cites.
OPTQ: Accurate quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Later among the works it cites.
A framework for few-shot language model evaluation, 2023
Gao, L., Tow, J., Abbasi, B., et al · 2023
Later among the works it cites.
llama.cpp, 2023
Gerganov, G · 2023
Later among the works it cites.
Repq-vit: Scale reparameterization for post-training quantization of vision transformers
Li, Z., Xiao, J., Yang, L., and Gu, Q · 2023
Later among the works it cites.
The cost of compression: Investigating the impact of compression on parametric knowledge in language models
Namburi, S. S. S., Sreedhar, M., Srinivasan, S., et al · 2023
Later among the works it cites.
A comprehensive survey on model quantization for deep neural networks in image classification
Rokh, B., Azarpeyvand, A., and Khanteymoori, A · 2023
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model, 2023
Scao, T. L., Fan, A., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., et al · 2023
Later among the works it cites.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Later among the works it cites.
Lookupffn: making transformers compute-lite for cpu inference
Zeng, Z., Davies, M., Pulijala, P., et al · 2023
Later among the works it cites.
Prompting large language model for machine translation: A case study
Zhang, B., Haddow, B., and Birch, A · 2023
Later among the works it cites.