Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have revolutionized natural language processing tasks.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 1905
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
Li, F.; Liu, B.; Wang, X.; Zhang, B.; and Yan, J. 2016 · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; McDonell, K.; Muennighoff, N.; et al. 2021 · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
PeQA: A Massive Persian Question-Answering and Chatbot Dataset
Arshia, F. Z.; Keyvanrad, M. A.; Sadidpour, S. S.; and Mohammadi, S. M. R. 2022 · 2022
Earlier work this paper cites.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Dettmers, T.; Lewis, M.; Belkada, Y.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z.; Yazdani Aminabadi, R.; Zhang, M.; Wu, X.; Li, C.; and He, Y. 2022 · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023 · 2023
Cited alongside, same era.
Spqr: A sparse-quantized representation for near-lossless llm weight compression
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G.; Lin, J.; Seznec, M.; Wu, H.; Demouth, J.; and Han, S. 2023 · 2023
Later among the works it cites.
Quarot: Outlier-free 4-bit inference in rotated llms
Ashkboos, S.; Mohtashami, A.; Croci, M. L.; Li, B.; Jaggi, M.; Alistarh, D.; Hoefler, T.; and Hensman, J. 2024 · 2024
Closest in time.
Low-Rank Quantization-Aware Training for LLMs
Bondarenko, Y.; Del Chiaro, R.; and Nagel, M. 2024 · 2024
Closest in time.
Quip: 2-bit quantization of large language models with guarantees
Chee, J.; Cai, Y.; Kuleshov, V.; and De Sa, C. M. 2024 · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Dettmers, T.; Pagnoni, A.; Holtzman, A.; and Zettlemoyer, L. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dettmers, T.; Svirschevski, R.; Egiazarian, V.; Kuznedelev, D.; Frantar, E.; Ashkboos, S.; Borzunov, A.; Hoefler, T.; and Alistarh, D. 2023 · 2023
Cited alongside, same era.
Large language models meet cognitive science: LLMs as tools, models, and participants
Hardy, M.; Sucholutsky, I.; Thompson, B.; and Griffiths, T. 2023 · 2023
Cited alongside, same era.
Unlocking the potential of user feedback: Leveraging large language model as user simulators to enhance dialogue system
Hu, Z.; Feng, Y.; Luu, A. T.; Hooi, B.; and Lipani, A. 2023 · 2023
Cited alongside, same era.
Squeezellm: Dense-and-sparse quantization
Kim, S.; Hooper, C.; Gholami, A.; Dong, Z.; Li, X.; Shen, S.; Mahoney, M. W.; and Keutzer, K. 2023 · 2023
Cited alongside, same era.
Llm-qat: Data-free quantization aware training for large language models
Liu, Z.; Oguz, B.; Zhao, C.; Chang, E.; Stock, P.; Mehdad, Y.; Shi, Y.; Krishnamoorthi, R.; and Chandra, V. 2023 · 2023
Cited alongside, same era.
Pb-llm: Partially binarized large language models
Shang, Y.; Yuan, Z.; Wu, Q.; and Dong, Z. 2023 · 2023
Cited alongside, same era.
Omniquant: Omnidirectionally calibrated quantization for large language models
Shao, W.; Chen, M.; Zhang, Z.; Xu, P.; Zhao, L.; Li, Z.; Zhang, K.; Gao, P.; Qiao, Y.; and Luo, P. 2023 · 2023
Cited alongside, same era.
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
Hu, X.; Chen, Y.; Yang, D.; Zhou, S.; Yuan, Z.; Yu, J.; and Xu, C. 2024a
Cited in the paper.
Egiazarian, V.; Panferov, A.; Kuznedelev, D.; Frantar, E.; Babenko, A.; and Alistarh, D. 2024 · 2024
Closest in time.
Billm: Pushing the limit of post-training quantization for llms
Huang, W.; Liu, Y.; Qin, H.; Li, Y.; Zhang, S.; Liu, X.; Magno, M.; and Qi, X. 2024 · 2024
Closest in time.
Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models
Lee, C.; Jin, J.; Kim, T.; Kim, H.; and Park, E. 2024 · 2024
Closest in time.
Quip#: Even better LLM quantization with hadamard incoherence and lattice codebooks
Tseng, A.; Chee, J.; Sun, Q.; Kuleshov, V.; and De Sa, C. 2024 · 2024
Closest in time.
Atom: Low-bit quantization for efficient and accurate llm serving
Zhao, Y.; Lin, C.-Y.; Zhu, K.; Ye, Z.; Chen, L.; Zheng, S.; Ceze, L.; Krishnamurthy, A.; Chen, T.; and Kasikci, B. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2024 · 2024
Closest in time.