Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) stand out for their impressive performance in intricate language modeling tasks.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020a · 1901
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
A Deep Look into Logarithmic Quantization of Model Parameters in Neural Networks
Cai, J.; Takemoto, M.; and Nakajo, H. 2018 · 2018
Earlier work this paper cites.
QNNPACK: Open source library for optimized mobile deep learning
Dukhan, M.; Wu, Y.; and Lu, H. 2018 · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B.; Kligys, S.; Chen, B.; Zhu, M.; Tang, M.; Howard, A.; Adam, H.; and Kalenichenko, D. 2018 · 2018
Earlier work this paper cites.
gemmlowp: A small self-contained low-precision gemm library
Jacob, B.; and Warden, P. 2017 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T.; Lewis, M.; Belkada, Y.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
Learned token pruning for transformers
Kim, S.; Shen, S.; Thorsley, D.; Gholami, A.; Kwon, W.; Hassoun, J.; and Keutzer, K. 2022 · 2022
Cited alongside, same era.
Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training
Kong, Z.; Ma, H.; Yuan, G.; Sun, M.; Xie, Y.; Dong, P.; Meng, X.; Shen, X.; Tang, H.; Qin, M.; et al. 2022 · 2022
Cited alongside, same era.
Not all patches are what you need: Expediting vision transformers via token reorganizations
Liang, Y.; Ge, C.; Tong, Z.; Song, Y.; Wang, J.; and Xie, P. 2022 · 2022
Cited alongside, same era.
A collection of low-level machine learning functions optimized with SIMD technologies
ARM. 2023 · 2023
Closest in time.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Dettmers, T.; Svirschevski, R.; Egiazarian, V.; Kuznedelev, D.; Frantar, E.; Ashkboos, S.; Borzunov, A.; Hoefler, T.; and Alistarh, D. 2023 · 2023
Closest in time.
Heatvit: Hardware-efficient adaptive token pruning for vision transformers
Dong, P.; Sun, M.; Lu, A.; Xie, Y.; Liu, K.; Kong, Z.; Meng, X.; Li, Z.; Lin, X.; Fang, Z.; et al. 2023 · 2023
Closest in time.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Dang, X.; and Han, S. 2023 · 2023
Closest in time.
Data Level Lottery Ticket Hypothesis for Vision Transformers
Shen, X.; Kong, Z.; Qin, M.; Dong, P.; Yuan, G.; Meng, X.; Tang, H.; Ma, X.; and Wang, Y. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer
Lin, Y.; Zhang, T.; Sun, P.; Li, Z.; and Zhou, S. 2022 · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilić, S.; Hesslow, D.; Castagné, R.; Luccioni, A. S.; Yvon, F.; Gallé, M.; et al. 2022 · 2022
Cited alongside, same era.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Xiao, G.; Lin, J.; Seznec, M.; Wu, H.; Demouth, J.; and Han, S. 2022 · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Cited alongside, same era.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020b
Cited in the paper.
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023 · 2023
Closest in time.
ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats
Wu, X.; Yao, Z.; and He, Y. 2023 · 2023
Closest in time.
Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models
Zhang, Y.; Zhao, L.; Cao, S.; Wang, W.; Cao, T.; Yang, F.; Yang, M.; Zhang, S.; and Xu, N. 2023 · 2023
Closest in time.