Fetching the paper…
Reading the bibliography…
Efficient adaption of large language models (LLMs) on edge devices is essential for applications requiring continuous and privacy-preserving adaptation and inference.
Pointer sentinel mixture models
Merity et al · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks. In ICPR
Teerapittayanon et al · 2016
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks et al · 2020
Earlier work this paper cites.
NVIDIA Jetson TX2
NVIDIA. 2020 · 2020
Earlier work this paper cites.
Enabling random precision switch for winning both adversarial robustness and efficiency. In MICRO . 225–237
Fu et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu et al · 2021
Earlier work this paper cites.
Understanding softmax confidence and uncertainty
Pearce et al · 2021
Earlier work this paper cites.
Unified visual transformer compression
Yu et al. 2022 · 2022
Cited alongside, same era.
Quest Pro
Meta. 2022 · 2022
Cited alongside, same era.
Lst: Ladder side-tuning for parameter and memory efficient transfer learning
Sung et al · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck et al · 2023
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Dettmers et al · 2023
Cited alongside, same era.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Frantar et al · 2023
Cited alongside, same era.
Systolic CNN AcceLErator Simulator (SCALE Sim)
Samajdar et al. 2023 · 2023
Later among the works it cites.
An Efficient Training Accelerator for Transformers With Hardware-Algorithm Co-Optimization
Shao et al · 2023
Later among the works it cites.
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Sheng et al · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron et al · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Zhang et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Liu et al · 2023
Cited alongside, same era.
Hint-aug: Drawing hints from foundation vision transformers towards boosted few-shot parameter-efficient tuning. In CVPR . 11102–11112
Yu et al. 2023a
Cited in the paper.
Master-ASR: achieving multilingual scalability and low-resource adaptation in ASR with modular learning. In ICML . PMLR, 40475–40487
Yu et al. 2023b
Cited in the paper.
Later among the works it cites.
Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization
Kim et al. 2024 · 2024
Closest in time.