Fetching the paper…
Reading the bibliography…
Large language models have been widely adopted but require significant GPU memory for inference.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S. (2019) · 1902
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. (2019) · 1904
Earlier work this paper cites.
Mixed precision training with 8-bit floating point
Mellempudi, N., Srinivasan, S., Das, D., and Kaul, B. (2019) · 1905
Earlier work this paper cites.
Representation degeneration problem in training natural language generation models
Gao, J., He, D., Tan, X., Qin, T., Wang, L., and Liu, T.-Y. (2019) · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2019) · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2019) · 1910
Earlier work this paper cites.
Training with quantization noise for extreme model compression
Fan, A., Stock, P., Graham, B., Grave, E., Gribonval, R., Jegou, H., and Joulin, A. (2020) · 2004
Earlier work this paper cites.
Binary neural networks: A survey
Qin, H., Gong, R., Liu, X., Bai, X., Song, J., and Sebe, N. (2020) · 2004
Earlier work this paper cites.
Integer quantization for deep learning inference: Principles and empirical evaluation
Wu, H., Judd, P., Zhang, X., Isaev, M., and Micikevicius, P. (2020) · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 2005
Earlier work this paper cites.
Towards fully 8-bit integer inference for the transformer model
Lin, Y., Li, Y., Liu, T., Xiao, T., Liu, T., and Zhu, J. (2020) · 2009
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al. (2020) · 2010
Earlier work this paper cites.
Binarybert: Pushing the limit of bert quantization
Bai, H., Zhang, W., Hou, L., Shang, L., Jin, J., Jiang, X., Liu, Q., Lyu, M. R., and King, I. (2021) · 2012
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Courbariaux, M., Bengio, Y., and David, J.-P. (2014) · 2014
Earlier work this paper cites.
Results of the wmt14 metrics shared task
Macháček, M. and Bojar, O. (2014) · 2014
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J. (2015) · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. (2015) · 2015
Earlier work this paper cites.
Binarynet: Training deep neural networks with weights and activations constrained to +1 or -1
Courbariaux, M. and Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Rastegari, M., Ordonez, V., Redmon, J., and Farhadi, A. (2016) · 2016
Earlier work this paper cites.
Edinburgh neural machine translation systems for wmt 16
Sennrich, R., Haddow, B., and Birch, A. (2016) · 2016
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Trained ternary quantization
Zhu, C., Han, S., Mao, H., and Dally, W. J. (2017) · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
Scaling neural machine translation
Ott, M., Edunov, S., Grangier, D., and Auli, M. (2018) · 2018
Cited alongside, same era.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al. (2018) · 2018
Cited alongside, same era.
Ternarybert: Distillation-aware ultra-low bit bert
Zhang, W., Hou, L., Yin, Y., Shang, L., Chen, X., Jiang, X., and Liu, Q. (2020) · 2020
Later among the works it cites.
Efficient large scale language modeling with mixtures of experts
Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X. V., Du, J., Iyer, S., Pasunuru, R., et al. (2021) · 2021
Later among the works it cites.
Understanding and overcoming the challenges of efficient transformer quantization
Bondarenko, Y., Nagel, M., and Blankevoort, T. (2021) · 2021
Later among the works it cites.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. (2021) · 2021
Later among the works it cites.
A survey of quantization methods for efficient neural network inference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A simple method for commonsense reasoning
Trinh, T. H. and Le, Q. V. (2018) · 2018
Cited alongside, same era.
Training deep neural networks with 8-bit floating point numbers
Wang, N., Choi, J., Brand, D., Chen, C., and Gopalakrishnan, K. (2018) · 2018
Cited alongside, same era.
Lq-nets: Learned quantization for highly accurate and compact deep neural networks
Zhang, D., Yang, J., Ye, D., and Hua, G. (2018) · 2018
Cited alongside, same era.
Accurate and efficient 2-bit quantized neural networks
Choi, J., Venkataramani, S., Srinivasan, V., Gopalakrishnan, K., Wang, Z., and Chuang, P. (2019) · 2019
Cited alongside, same era.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K. (2019) · 2019
Cited alongside, same era.
Openwebtext corpus
Gokaslan, A. and Cohen, V. (2019) · 2019
Cited alongside, same era.
Differentiable soft quantization: Bridging full-precision and low-bit neural networks
Gong, R., Liu, X., Jiang, S., Li, T., Hu, P., Lin, J., Yu, F., and Yan, J. (2019) · 2019
Cited alongside, same era.
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K. (2021) · 2021
Later among the works it cites.
Fbgemm: Enabling high-performance low-precision deep learning inference
Khudia, D., Huang, J., Basu, P., Deng, S., Liu, H., Park, J., and Smelyanskiy, M. (2021) · 2021
Later among the works it cites.
Bert busters: Outlier dimensions that disrupt transformers
Kovaleva, O., Kulshreshtha, S., Rogers, A., and Rumshisky, A. (2021) · 2021
Later among the works it cites.
Positional artefacts propagate through masked language model embeddings
Luo, Z., Kulmizev, A., and Mao, X. (2021) · 2021
Later among the works it cites.
Timkey, W. and van Schijndel, M. (2021) · 2021
Later among the works it cites.
Hawq-v3: Dyadic neural network quantization
Yao, Z., Dong, Z., Zheng, Z., Gholami, A., Yu, J., Tan, E., Wang, L., Huang, Q., Wang, Y., Mahoney, M., et al. (2021) · 2021
Later among the works it cites.
Automatic mixed-precision quantization search of bert
Zhao, C., Hua, T., Shen, Y., Lou, Q., and Jin, H. (2021) · 2021
Later among the works it cites.
8-bit optimizers via block-wise quantization
Dettmers, T., Lewis, M., Shleifer, S., and Zettlemoyer, L. (2022) · 2022
Closest in time.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. (2022) · 2022
Closest in time.
F8net: Fixed-point 8-bit only multiplication for network quantization
Jin, Q., Ren, J., Zhuang, R., Hanumante, S., Li, Z., Chen, Z., Wang, Y., Yang, K., and Tulyakov, S. (2022) · 2022
Closest in time.
nuqmm: Quantized matmul for efficient inference of large-scale generative language models
Park, G., Park, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D. (2022) · 2022
Closest in time.
Outliers dimensions that disrupt transformers are driven by frequency
Puccetti, G., Rogers, A., Drozd, A., and Dell’Orletta, F. (2022) · 2022
Closest in time.
Outlier suppression: Pushing the limit of low-bit transformer language models
Wei, X., Zhang, Y., Zhang, X., Gong, R., Zhang, S., Zhang, Q., Yu, F., and Liu, X. (2022) · 2022
Closest in time.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y. (2022) · 2022
Closest in time.
Glm-130b: An open bilingual pre-trained model
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., et al. (2022) · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. (2022) · 2022
Closest in time.