Fetching the paper…
Reading the bibliography…
Transformer plays a vital role in the realms of natural language processing (NLP) and computer vision (CV), specially for constructing large language models (LLM) and large vision models (LVM).
Linear system theory and design
C.-T. Chen · 1984
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart et al · 1986
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Model compression
C. Bucilua et al · 2006
Earlier work this paper cites.
Early exit optimizations for additive machine learned ranking systems
B. B. Cambazoglu et al · 2010
Earlier work this paper cites.
Predicting parameters in deep learning
M. Denil et al · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero et al · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
M. Jaderberg et al · 2014
Earlier work this paper cites.
S. Han et al · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton et al · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Y. Kim and A. M. Rush · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
S. Teerapittayanon et al · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani et al · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang et al · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
M. Zhu and S. Gupta · 2017
Earlier work this paper cites.
Channel pruning for accelerating very deep neural networks
Y. He et al · 2017
Earlier work this paper cites.
Fast approximate computations with cauchy matrices and polynomials
V. Pan · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin et al · 2018
Earlier work this paper cites.
Deep contextualized word representations
M. E. Peters et al · 2018
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
N. Lee et al · 2018
Earlier work this paper cites.
A systematic dnn weight pruning framework using alternating direction method of multipliers
T. Zhang et al · 2018
Earlier work this paper cites.
Image transformer
N. Parmar et al · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford et al · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh et al · 2019
Earlier work this paper cites.
Q8bert: Quantized 8bit bert
O. Zafrir et al · 2019
Earlier work this paper cites.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Z. Dong et al · 2019
Earlier work this paper cites.
Distilling task-specific knowledge from bert into simple neural networks
R. Tang et al · 2019
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
I. Turc et al · 2019
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
S. Sun et al · 2019
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
X. Jiao et al · 2019
Earlier work this paper cites.
Hint-based training for non-autoregressive machine translation
Z. Li et al · 2019
Earlier work this paper cites.
Small and practical bert models for sequence labeling
H. Tsai et al · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin et al · 2019
Earlier work this paper cites.
Gate decorator: Global filter pruning method for accelerating deep convolutional neural networks
Z. You et al · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
P. Michel et al · 2019
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
A. Fan et al · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
E. Voita et al · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child et al · 2019
Earlier work this paper cites.
Q. Guo et al · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child et al · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
J. W. Rae et al · 2019
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
U. Khandelwal et al · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
N. Shazeer · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi et al · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke et al · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf et al · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai et al · 2019
Earlier work this paper cites.
Bert and pals: Projected attention layers for efficient adaptation in multi-task learning
A. C. Stickland and I. Murray · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy et al · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
W. Wang et al · 2020
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
H. Touvron et al · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
N. Kitaev et al · 2020
Earlier work this paper cites.
Q-bert: Hessian based ultra low precision quantization of bert
S. Shen et al · 2020
Earlier work this paper cites.
Binarybert: Pushing the limit of bert quantization
H. Bai et al · 2020
Earlier work this paper cites.
Ternarybert: Distillation-aware ultra-low bit bert
W. Zhang et al · 2020
Earlier work this paper cites.
Mixkd: Towards efficient distillation of large-scale language models
K. J. Liang et al · 2020
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Z. Sun et al · 2020
Earlier work this paper cites.
Xtremedistil: Multi-stage distillation for massive multilingual models
S. Mukherjee and A. Awadallah · 2020
Earlier work this paper cites.
Bert-of-theseus: Compressing bert by progressive module replacing
C. Xu et al · 2020
Earlier work this paper cites.
Pre-trained models for natural language processing: A survey
X. Qiu et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel et al · 2020
Earlier work this paper cites.
Comparing rewinding and fine-tuning in neural network pruning
A. Renda et al · 2020
Earlier work this paper cites.
Power-bert: Accelerating bert inference via progressive word-vector elimination
S. Goyal et al · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
M. Zaheer et al · 2020
Earlier work this paper cites.
Efficient content-based sparse attention with routing transformers
A. Roy et al · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion et al · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
M. Zaheer et al · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy et al · 2020
Earlier work this paper cites.
Faster transformer decoding: N-gram masked self-attention
C. Chelba et al · 2020
Earlier work this paper cites.
Rethinking attention with performers
K. Choromanski et al · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos et al · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley et al · 2020
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
A. Gu et al · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin et al · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
The right tool for the job: Matching model and instance complexities
R. Schwartz et al · 2020
Earlier work this paper cites.
Fastbert: a self-distilling bert with adaptive inference time
W. Liu et al · 2020
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
L. Hou et al · 2020
Earlier work this paper cites.
Depth-adaptive transformer
M. Elbayad et al · 2020
Earlier work this paper cites.
Apq: Joint search for network architecture, pruning and quantization policy
T. Wang et al · 2020
Earlier work this paper cites.
Pre-trained image processing transformer
H. Chen et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford et al · 2021
Earlier work this paper cites.
Post-training quantization for vision transformer
Z. Liu et al · 2021
Earlier work this paper cites.
Fq-vit: Post-training quantization for fully quantized vision transformer
Y. Lin et al · 2021
Earlier work this paper cites.
Patch slimming for efficient vision transformers
Y. Tang et al · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu et al · 2021
Earlier work this paper cites.
Mlp-mixer: An all-mlp architecture for vision
I. O. Tolstikhin et al · 2021
Earlier work this paper cites.
Understanding and overcoming the challenges of efficient transformer quantization
Y. Bondarenko et al · 2021
Earlier work this paper cites.
I-bert: Integer-only bert quantization
S. Kim et al · 2021
Earlier work this paper cites.
Automatic mixed-precision quantization search of bert
C. Zhao et al · 2021
Earlier work this paper cites.
Symbolic knowledge distillation: from general language models to commonsense models
P. West et al · 2021
Earlier work this paper cites.
A short study on compressing decoder-based language models
T. Li et al · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
H. Touvron et al · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu et al · 2021
Cited alongside, same era.
Learning n: M fine-grained structured sparse neural networks from scratch
A. Zhou et al · 2021
Cited alongside, same era.
Block pruning for faster transformers
F. Lagunas et al · 2021
Cited alongside, same era.
Accelerated sparse neural training: A provable and efficient method to find n: M transposable masks
I. Hubara et al · 2021
Cited alongside, same era.
Kvt: k-nn attention for boosting vision transformers
P. Wang et al · 2022
Later among the works it cites.
Fast vision transformers with hilo attention
Z. Pan et al · 2022
Later among the works it cites.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
X. Dong et al · 2022
Later among the works it cites.
Msg-transformer: Exchanging local spatial information by manipulating messenger tokens
J. Fang et al · 2022
Later among the works it cites.
Resmlp: Feedforward networks for image classification with data-efficient training
H. Touvron et al · 2022
Later among the works it cites.
Hire-mlp: Vision mlp via hierarchical rearrangement
J. Guo et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. E. Hu et al · 2021
Cited alongside, same era.
Learned token pruning for transformers
S. Kim et al · 2021
Cited alongside, same era.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Y. Rao et al · 2021
Cited alongside, same era.
Chasing sparsity in vision transformers: An end-to-end exploration
T. Chen et al · 2021
Cited alongside, same era.
Vision transformer pruning
M. Zhu et al · 2021
Cited alongside, same era.
Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition
Y. Wang et al · 2021
Cited alongside, same era.
Sparse detr: Efficient end-to-end object detection with learnable sparsity
B. Roh et al · 2021
Cited alongside, same era.
An image patch is a wave: Phase-aware vision mlp
Y. Tang et al · 2022
Later among the works it cites.
Scaling vision transformers
X. Zhai et al · 2022
Later among the works it cites.
Vitas: Vision transformer architecture search
X. Su et al · 2022
Later among the works it cites.
Kronecker decomposition for gpt compression
A. Edalati et al · 2022
Later among the works it cites.
Confident adaptive language modeling
T. Schuster et al · 2022
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron et al · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Lion: Adversarial distillation of closed-source large language model
Y. Jiang et al · 2023
Later among the works it cites.
Sequential modeling enables scalable learning for large vision models, 2023
Y. Bai et al · 2023
Later among the works it cites.
H. Liu et al · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
G. Xiao et al · 2023
Later among the works it cites.
Omniquant: Omnidirectionally calibrated quantization for large language models
W. Shao et al · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
T. Dettmers et al · 2023
Later among the works it cites.
Oscillation-free quantization for low-bit vision transformers
S.-Y. Liu et al · 2023
Later among the works it cites.
Llm-pruner: On the structural pruning of large language models
X. Ma et al · 2023
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
M. Xia et al · 2023
Later among the works it cites.
Dynamic context pruning for efficient and interpretable autoregressive transformers
S. Anagnostidis et al · 2023
Later among the works it cites.
X-pruner: explainable pruning for vision transformers
L. Yu and W. Xiang · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Y. Sun et al · 2023
Later among the works it cites.
Noisyquant: Noisy bias-enhanced post-training activation quantization for vision transformers
Y. Liu et al · 2023
Later among the works it cites.
Repq-vit: Scale reparameterization for post-training quantization of vision transformers
Z. Li et al · 2023
Later among the works it cites.
Data-free quantization via mixed-precision compensation without fine-tuning
J. Chen et al · 2023
Later among the works it cites.
I-vit: Integer-only quantization for efficient vision transformer inference
Z. Li and Q. Gu · 2023
Later among the works it cites.
Jumping through local minima: Quantization in the loss landscape of vision transformers
N. Frumkin et al · 2023
Later among the works it cites.
Patch-wise mixed-precision quantization of vision transformer
J. Xiao et al · 2023
Later among the works it cites.
Psaq-vit v2: Toward accurate and general data-free quantization for vision transformers
Z. Li et al · 2023
Later among the works it cites.
Q-detr: An efficient low-bit quantized detection transformer
S. Xu et al · 2023
Later among the works it cites.
Variation-aware vision transformer quantization
X. Huang et al · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
J. Lin et al · 2023
Later among the works it cites.
X. Wei et al · 2023
Later among the works it cites.
Qllm: Accurate and efficient low-bitwidth quantization for large language models
J. Liu et al · 2023
Later among the works it cites.
Rptq: Reorder-based post-training quantization for large language models
Z. Yuan et al · 2023
Later among the works it cites.
Cbq: Cross-block quantization for large language models
X. Ding et al · 2023
Later among the works it cites.
Squeezellm: Dense-and-sparse quantization
S. Kim et al · 2023
Later among the works it cites.
Optimize weight rounding via signed gradient descent for the quantization of llms
W. Cheng et al · 2023
Later among the works it cites.
Llm-qat: Data-free quantization aware training for large language models
Z. Liu et al · 2023
Later among the works it cites.
Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization
J. Kim et al · 2023
Later among the works it cites.
Weight-inherited distillation for task-agnostic bert compression
T. Wu et al · 2023
Later among the works it cites.
Knowledge distillation of large language models
Y. Gu et al · 2023
Later among the works it cites.
Gkd: Generalized knowledge distillation for auto-regressive sequence models
R. Agarwal et al · 2023
Later among the works it cites.
S. Bae et al · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
R. Schaeffer et al · 2023
Later among the works it cites.
C.-Y. Hsieh et al · 2023
Later among the works it cites.
Scott: Self-consistent chain-of-thought distillation
P. Wang et al · 2023
Later among the works it cites.
Can language models teach weaker agents? teacher explanations improve students via theory of mind
S. Saha et al · 2023
Later among the works it cites.
Lamini-lm: A diverse herd of distilled models from large-scale instructions
M. Wu et al · 2023
Later among the works it cites.
Pad: Program-aided distillation specializes large models in reasoning
X. Zhu et al · 2023
Later among the works it cites.
Distilling reasoning capabilities into smaller language models
K. Shridhar et al · 2023
Later among the works it cites.
Specializing smaller language models towards multi-step reasoning
Y. Fu et al · 2023
Later among the works it cites.
Large language model distillation doesn’t need a teacher
A. H. Jha et al · 2023
Later among the works it cites.
Vanillakd: Revisit the power of vanilla knowledge distillation from small scale to large scale
Z. Hao et al · 2023
Later among the works it cites.
Cumulative spatial knowledge distillation for vision transformers
B. Zhao et al · 2023
Later among the works it cites.
A survey on deep neural network pruning-taxonomy, comparison, analysis, and recommendations
H. Cheng et al · 2023
Later among the works it cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
E. Frantar and D. Alistarh · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
M. Sun et al · 2023
Later among the works it cites.
Prune and tune: Improving efficient pruning techniques for massive language models
A. Syed et al · 2023
Later among the works it cites.
What matters in the structured pruning of generative language models?
M. Santacroce et al · 2023
Later among the works it cites.
Structured pruning for efficient generative pre-trained language models
C. Tao et al · 2023
Later among the works it cites.
Sparse token transformer with attention back tracking
H. Lee et al · 2023
Later among the works it cites.
A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking
L. Papa et al · 2023
Later among the works it cites.
Less is more: Focus attention for efficient detr
D. Zheng et al · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2023
Later among the works it cites.
G. Penedo et al · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
G. Xiao et al · 2023
Later among the works it cites.
Starcoder: may the source be with you!
R. Li et al · 2023
Later among the works it cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
J. Ainslie et al · 2023
Later among the works it cites.
Colossal-ai: A unified deep learning system for large-scale parallel training
S. Li et al · 2023
Later among the works it cites.
Hyena hierarchy: Towards larger convolutional language models
M. Poli et al · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
B. Peng et al · 2023
Later among the works it cites.
Towards a unified view of sparse feed-forward network in pretraining large language model
L. Z. Liu et al · 2023
Later among the works it cites.
Fastervit: Fast vision transformers with hierarchical attention
A. Hatamizadeh et al · 2023
Later among the works it cites.
Efficientvit: Memory efficient vision transformer with cascaded group attention
X. Liu et al · 2023
Later among the works it cites.
Riformer: Keep your vision backbone effective but removing token mixer
J. Wang et al · 2023
Later among the works it cites.
Flatten transformer: Vision transformer using focused linear attention
D. Han et al · 2023
Later among the works it cites.
Vanillanet: the power of minimalism in deep learning
H. Chen et al · 2023
Later among the works it cites.
Lord: Low rank decomposition of monolingual code llms for one-shot compression
A. Kaushal et al · 2023
Later among the works it cites.
Ternary singular value decomposition as a better parameterized form in linear mapping
B. Chen et al · 2023
Later among the works it cites.
M. Xu et al · 2023
Later among the works it cites.
Skipdecode: Autoregressive skip decoding with batching and caching for efficient llm inference
L. Del Corro et al · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Y. Leviathan et al · 2023
Later among the works it cites.
Accelerating large language model decoding with speculative sampling
C. Chen et al · 2023
Later among the works it cites.
Inference with reference: Lossless acceleration of large language models
N. Yang et al · 2023
Later among the works it cites.
Llmcad: Fast and scalable on-device large language model inference
D. Xu et al · 2023
Later among the works it cites.
Rethinking kullback-leibler divergence in knowledge distillation for large language models
T. Wu et al · 2024
Closest in time.
Distillm: Towards streamlined distillation for large language models
J. Ko et al · 2024
Closest in time.
Moe-mamba: Efficient selective state space models with mixture of experts
M. Pióro et al · 2024
Closest in time.
Densemamba: State space models with dense hidden connection for efficient large language models
W. He et al · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
L. Zhu et al · 2024
Closest in time.
Vmamba: Visual state space model
Y. Liu et al · 2024
Closest in time.