Fetching the paper…
Reading the bibliography…
Key-value (KV) caching has become the de-facto to accelerate generation speed for large language models (LLMs) inference.
Powersgd: Practical low-rank gradient compression for distributed optimization
Vogels, T · 1905
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T · 1910
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Tenney, I · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. https://aclanthology.org/P19-1580
Voita, E · 2019
Earlier work this paper cites.
Q8bert: Quantized 8bit bert
Zafrir, O · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K · 2021
Earlier work this paper cites.
Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T · 2022
Earlier work this paper cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T · 2022
Earlier work this paper cites.
Efficiently scaling transformer inference
Pope, R · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M · 2022
Cited alongside, same era.
Lamda: Language models for dialog applications
Thoppilan, R · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J · 2022
Cited alongside, same era.
Wordcraft: Story writing with large language models
Yuan, A · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S · 2022
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y · 2023
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time
Liu, Z · 2023
Later among the works it cites.
Relu strikes back: Exploiting activation sparsity in large language models
Mirzadeh, I · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Sparq attention: Bandwidth-efficient llm inference
Ribar, L · 2023
Later among the works it cites.
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng, Y · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frantar, E · 2023
Cited alongside, same era.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Fu, Y · 2023
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms
Ge, S · 2023
Cited alongside, same era.
Mistral 7b
Jiang, A. Q · 2023
Cited alongside, same era.
Squeezellm: Dense-and-sparse quantization
Kim, S · 2023
Cited alongside, same era.
Loftq: Lora-fine-tuning-aware quantization for large language models
Li, Y · 2023
Cited alongside, same era.
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G · 2023
Later among the works it cites.
Atom: Low-bit quantization for efficient and accurate llm serving
Zhao, Y · 2023
Later among the works it cites.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C · 2024
Closest in time.
Kivi: A tuning-free asymmetric 2bit quantization for kv cache
Liu, Z · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date. https://ai.meta.com/blog/meta-llama-3/
Meta · 2024
Closest in time.
Tell your model where to attend: Post-hoc attention steering for LLMs
Zhang, Q · 2024
Closest in time.