Fetching the paper…
Reading the bibliography…
Multi-modal Large Language Models (MLLMs) serving systems commonly employ KV-cache compression to reduce memory footprint.
Lynx: Using os and hardware support for fast fine-grained inter-core communication
2016
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
2019
Earlier work this paper cites.
Memory-efficient pipeline-parallel dnn training
2021
Earlier work this paper cites.
Towards understanding the mixture-of-experts layer in deep learning
2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
2022
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
2022
Earlier work this paper cites.
Mixture-of-experts with expert choice routing
2022
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
2023
Earlier work this paper cites.
Model tells you what to discard: Adaptive kv cache compression for llms
2023
Earlier work this paper cites.
Compressed context memory for online language model interaction
2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
2023
Earlier work this paper cites.
Pond: Cxl-based memory pooling systems for cloud platforms
2023
Earlier work this paper cites.
Visual instruction tuning, 2023
2023
Earlier work this paper cites.
Gpt-4 technical report. arxiv 2303.08774
2023
Cited alongside, same era.
Efficiently scaling transformer inference
2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
2023
Cited alongside, same era.
Keyformer: Kv cache reduction through key tokens selection for efficient generative inference
2024
Cited alongside, same era.
Snapkv: Llm knows what you are looking for before generation
2024
Later among the works it cites.
Cachegen: Kv cache compression and streaming for fast large language model serving
2024
Later among the works it cites.
Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
2024
Later among the works it cites.
Kivi: A tuning-free asymmetric 2bit quantization for kv cache
2024
Later among the works it cites.
Dynamic memory compression: Retrofitting llms for accelerated inference
2024
Later among the works it cites.
Kvpress: An nvidia hardware-accelerated framework for key-value cache compression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimised grouped-query attention mechanism for transformers
2024
Cited alongside, same era.
Weighted grouped query attention in transformers
2024
Cited alongside, same era.
Starnuma: Mitigating numa challenges with memory pooling
2024
Cited alongside, same era.
Apparate: Rethinking early exits to tame latency-throughput tensions in ml serving
2024
Cited alongside, same era.
A simple and effective
2024
Cited alongside, same era.
Milebench: Benchmarking mllms in long context
2024
Cited alongside, same era.
2024
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting
2024
Later among the works it cites.
Mooncake: A kvcache-centric disaggregated architecture for llm serving
2024
Later among the works it cites.
Look-m: Look-once optimization in kv cache for efficient multimodal long-context inference
2024
Later among the works it cites.
Model tells you where to merge: Adaptive kv cache merging for llms on long-context tasks
2024
Later among the works it cites.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
2024
Later among the works it cites.
vtensor: Flexible virtual tensor management for efficient llm serving
2024
Later among the works it cites.
Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling
2024
Later among the works it cites.
Alisa: Accelerating large language model inference via sparsity-aware kv caching
2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs
2024
Later among the works it cites.