Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated impressive abilities in various domains while the inference cost is expensive.
Position-location solutions by Taylor-series estimation
Wade H Foy. 1976 · 1976
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, et al · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, et al · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language. In AAAI
Yonatan Bisk, Rowan Zellers, et al · 2020
Earlier work this paper cites.
Sign language transformers: Joint end-to-end sign language recognition and translation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10023–10033
Necati Cihan Camgoz, Oscar Koller, Simon Hadfield, and Richard Bowden. 2020 · 2020
Earlier work this paper cites.
Q-bert: Hessian based ultra low precision quantization of bert. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 8815–8821
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020 · 2020
Earlier work this paper cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 811–824
Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. 2020 · 2020
Earlier work this paper cites.
A white paper on neural network quantization
Markus Nagel, Marios Fournarakis, et al · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, et al · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, et al · 2022
Earlier work this paper cites.
Shortcut learning of large language models in natural language understanding: A survey
Mengnan Du et al · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022 · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Glm-130b: An open bilingual pre-trained model
Aohan Zeng, Xiao Liu, et al · 2022
Cited alongside, same era.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Lever: Learning to verify language-to-code generation with execution. In International Conference on Machine Learning
Ansong Ni, Srini Iyer, et al · 2023
Closest in time.
Baolin Peng, Chunyuan Li, et al · 2023
Closest in time.
Omniquant: Omnidirectionally calibrated quantization for large language models
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, et al · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models. In ICML
Guangxuan Xiao, Ji Lin, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tim Dettmers et al · 2023
Cited alongside, same era.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar and Dan Alistarh. 2023 · 2023
Cited alongside, same era.
Advanced Ultra-Low Bitrate Compression Techniques for the LLaMA Family of LLMs
Nianhui Guo et al · 2023
Cited alongside, same era.
SqueezeLLM: Dense-and-Sparse Quantization
Sehoon Kim, Coleman Hooper, et al · 2023
Cited alongside, same era.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, et al · 2023
Cited alongside, same era.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Zechun Liu, Barlas Oguz, et al · 2023
Cited alongside, same era.
ChatGPT and a new academic reality: Artificial Intelligence-written research papers and the ethics of the large language models in scholarly publishing
Brady D Lund, Ting Wang, Nishith Reddy Mannuru, Bing Nie, Somipam Shimray, and Ziang Wang. 2023 · 2023
Cited alongside, same era.
Introducing LLaMA: A foundational, 65-billion-parameter large language model
AI Meta. 2023 · 2023
Cited alongside, same era.
Closest in time.
Average Price: Electricity per Kilowatt-Hour in U.S. City Average
FRED. 2024 · 2024
Closest in time.
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
Ziyi Guan, Hantao Huang, Yupeng Su, Hong Huang, Ngai Wong, and Hao Yu. 2024 · 2024
Closest in time.
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, and Xiaojuan Qi. 2024 · 2024
Closest in time.
HuggingFace. 2024
2024
Closest in time.
Abstractive Long Text Summarization Using Large Language Models
Gunjan Keswani, Wani Bisen, Hirkani Padwad, Yash Wankhedkar, Sudhanshu Pandey, and Ayushi Soni. 2024 · 2024
Closest in time.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2024 · 2024
Closest in time.
Build the future of AI with Meta Llama 3
Meta. 2024 · 2024
Closest in time.
The Inference Cost Of Search Disruption - Large Language Model Cost Analysis
Dylan Patel and Afzal Ahmad. 2024 · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. 2024 · 2024
Closest in time.