Fetching the paper…
Reading the bibliography…
Large language models (LLMs) demonstrate strong performance across natural language processing tasks, yet undergo significant performance degradation when modified for deployment through quantization, pruning, or decoding strategy adjustments.
Scaling laws for neural language models (2020)
Kaplan, J. et al · 2001
Earlier work this paper cites.
Understanding the difficulty of training transformers (2020)
Liu, L., Liu, X., Gao, J., Chen, W. & Han, J · 2004
Earlier work this paper cites.
Layer normalization (2016)
Ba, J. L., Kiros, J. R. & Hinton, G. E · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A. et al · 2017
Earlier work this paper cites.
Beam search strategies for neural machine translation
Freitag, M. & Al-Onaizan, Y · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N. et al · 2017
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B. et al · 2018
Earlier work this paper cites.
Diverse beam search for improved description of complex scenes
Vijayakumar, A. et al · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M. & Dauphin, Y · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N. & Uszkoreit, J · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. & Carbin, M · 2019
Earlier work this paper cites.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
Morcos, A., Yu, H., Paganini, M. & Tian, Y · 2019
Earlier work this paper cites.
Haq: Hardware-aware automated quantization with mixed precision
Wang, K., Liu, Z., Lin, Y., Lin, J. & Han, S · 2019
Earlier work this paper cites.
Stochastic beams and where to find them: The gumbel-top-k trick for sampling sequences without replacement
Kool, W., Van Hoof, H. & Welling, M · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. & Sennrich, R · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C. et al · 2020
Earlier work this paper cites.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R. & Modha, D. S · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M. & Choi, Y · 2020
Earlier work this paper cites.
On layer normalization in the transformer architecture
Xiong, R. et al · 2020
Earlier work this paper cites.
Improving transformer optimization through better initialization
Huang, X. S., Perez, F., Ba, J. & Volkovs, M · 2020
Earlier work this paper cites.
Cat-gen: Improving robustness in nlp models via controlled adversarial text generation
Wang, T. et al · 2020
Earlier work this paper cites.
Pretrained transformers improve out-of-distribution robustness
Hendrycks, D. et al · 2020
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A. et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A. et al · 2021
Earlier work this paper cites.
Carbon emissions and large neural network training (2021)
Patterson, D. et al · 2021
Earlier work this paper cites.
Mirostat: A neural text decoding algorithm that directly controls perplexity
Basu, S., Ramachandran, G. S., Keskar, N. S. & Varshney, L. R · 2021
Earlier work this paper cites.
Evaluating large language models trained on code (2021)
Chen, M. et al · 2021
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ding, M. et al · 2021
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D. et al · 2021
Earlier work this paper cites.
M6-t: Exploring sparse expert models and beyond (2021)
Yang, A. et al · 2021
Earlier work this paper cites.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N. & Zettlemoyer, L · 2021
Earlier work this paper cites.
Understanding robustness of transformers for image classification
Bhojanapalli, S. et al · 2021
Earlier work this paper cites.
On the robustness of vision transformers to adversarial examples
Mahmood, K., Mahmood, R. & Van Dijk, M · 2021
Earlier work this paper cites.
Reveal of vision transformers robustness against adversarial attacks (2021)
Aldahdooh, A., Hamidouche, W. & Deforges, O · 2021
Earlier work this paper cites.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Wang, B. et al · 2021
Earlier work this paper cites.
Robustbench: a standardized adversarial robustness benchmark
Croce, F. et al · 2021
Earlier work this paper cites.
Towards robustness against natural language word substitutions
Dong, X., Luu, A. T., Ji, R. & Liu, H · 2021
Earlier work this paper cites.
Robustness gym: Unifying the nlp evaluation landscape
Goel, K. et al · 2021
Earlier work this paper cites.
Evaluating the robustness of neural language models to input perturbations
Moradi, M. & Samwald, M · 2021
Earlier work this paper cites.
Training compute-optimal large language models (2022)
Hoffmann, J. et al · 2022
Earlier work this paper cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z. et al · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models (2022)
Zhang, S. et al · 2022
Earlier work this paper cites.
A contrastive framework for neural text generation
Su, Y. et al · 2022
Earlier work this paper cites.
On layer normalizations and residual connections in transformers (2022)
Takase, S., Kiyono, S., Kobayashi, S. & Suzuki, J · 2022
Earlier work this paper cites.
Foundation transformers (2022)
Wang, H. et al · 2022
Cited alongside, same era.
Glm-130b: An open bilingual pre-trained model (2022)
Zeng, A. et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B. & Shazeer, N · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N. et al · 2022
Cited alongside, same era.
Mixture-of-experts with expert choice routing
Zhou, Y. et al · 2022
Cited alongside, same era.
Stablemoe: Stable routing strategy for mixture of experts
Dai, D. et al · 2022
Cited alongside, same era.
Explorations of self-repair in language models
Rushing, C. & Nanda, N · 2024
Later among the works it cites.
The remarkable robustness of llms: Stages of inference? (2024)
Lad, V., Gurnee, W. & Tegmark, M · 2024
Later among the works it cites.
The unreasonable ineffectiveness of the deeper layers
Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P. & Roberts, D · 2024
Later among the works it cites.
A survey on model compression for large language models
Zhu, X., Li, J., Liu, Y., Ma, C. & Wang, W · 2024
Later among the works it cites.
Fluctuation-based adaptive structured pruning for large language models
An, Y., Zhao, X., Yu, T., Tang, M. & Wang, J · 2024
Later among the works it cites.
Shortened llama: Depth pruning for large language models with comparison of retraining methods (2024)
Kim, B.-K. et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paul, S. & Chen, P.-Y · 2022
Cited alongside, same era.
Towards robust vision transformer
Mao, X. et al · 2022
Cited alongside, same era.
Assaying out-of-distribution generalization in transfer learning
Wenzel, F. et al · 2022
Cited alongside, same era.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman, R., Logan IV, R., Matt, G. & Sameer, S · 2022
Cited alongside, same era.
Measure and improve robustness in nlp models: A survey
Wang, X., Wang, H. & Yang, D · 2022
Cited alongside, same era.
The king is naked: On the notion of robustness for natural language processing
La Malfa, E. & Kwiatkowska, M · 2022
Cited alongside, same era.
Slicegpt: Compress large language models by deleting rows and columns (2024)
Ashkboos, S., Croci, M. L., Nascimento, M. G. d., Hoefler, T. & Hensman, J · 2024
Later among the works it cites.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A. & Kolter, J. Z · 2024
Later among the works it cites.
Plug-and-play: An efficient post-training pruning method for large language models
Zhang, Y. et al · 2024
Later among the works it cites.
Fast and effective weight update for pruned large language models
Boža, V · 2024
Later among the works it cites.
Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity
Yin, L. et al · 2024
Later among the works it cites.
Besa: Pruning large language models with blockwise parameter-efficient sparsity allocation
Xu, P. et al · 2024
Later among the works it cites.
Dynamic sparse no training: Training-free fine-tuning for sparse llms
Zhang, Y. et al · 2024
Later among the works it cites.
Llm-qat: Data-free quantization aware training for large language models
Liu, Z. et al · 2024
Later among the works it cites.
Contemporary advances in neural network quantization: A survey
Li, M. et al · 2024
Later among the works it cites.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Lin, J. et al · 2024
Later among the works it cites.
Break the sequential dependency of llm inference using lookahead decoding
Fu, Y., Bailis, P., Stoica, I. & Zhang, H · 2024
Later among the works it cites.
Eagle: Speculative sampling requires rethinking feature uncertainty
Li, Y., Wei, F., Zhang, C. & Zhang, H · 2024
Later among the works it cites.
A thorough examination of decoding methods in the era of llms
Shi, C. et al · 2024
Later among the works it cites.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T. et al · 2024
Later among the works it cites.
A frustratingly simple decoding method for neural text generation
Yang, H. et al · 2024
Later among the works it cites.
Deepseek-v3 technical report (2024)
Liu, A. et al · 2024
Later among the works it cites.
Qwen2. 5 technical report (2024)
Yang, A. et al · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size (2024)
Team, G. et al · 2024
Later among the works it cites.
Deepnet: Scaling transformers to 1,000 layers
Wang, H. et al · 2024
Later among the works it cites.
Performance law of large language models (2024)
Wu, C. & Tang, R · 2024
Later among the works it cites.
Poisonbench: Assessing large language model vulnerability to data poisoning (2024)
Fu, T. et al · 2024
Later among the works it cites.
A survey on mixture of experts (2024)
Cai, W. et al · 2024
Later among the works it cites.
Yuan 2.0-m32: Mixture of experts with attention router (2024)
Wu, S. et al · 2024
Later among the works it cites.
Turn waste into worth: Rectifying top- k k router of moe (2024)
Zeng, Z. et al · 2024
Later among the works it cites.
Benchmarking large multimodal models against common corruptions (2024)
Zhang, J. et al · 2024
Later among the works it cites.
Large language models sensitivity to the order of options in multiple-choice questions
Pezeshkpour, P. & Hruschka, E · 2024
Later among the works it cites.
Lost in the middle: How language models use long contexts
Liu, N. F. et al · 2024
Later among the works it cites.
∞ \infty bench: Extending long context evaluation beyond 100k tokens
Zhang, X. et al · 2024
Later among the works it cites.
RULER: What’s the real context size of your long-context language models?
Hsieh, C.-P. et al · 2024
Later among the works it cites.
Long context rag performance of large language models (2024)
Leng, Q., Portes, J., Havens, S., Zaharia, M. & Carbin, M · 2024
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. & Dao, T · 2024
Later among the works it cites.
Jamba: A hybrid transformer-mamba language model (2024)
Lieber, O. et al · 2024
Later among the works it cites.
Understanding and mitigating the label noise in pre-training on downstream tasks
Chen, H. et al · 2024
Later among the works it cites.
The llama 3 herd of models (2024)
Dubey, A. et al · 2024
Later among the works it cites.
Yi: Open foundation models by 01. ai (2024)
Young, A. et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning (2025)
Guo, D. et al · 2025
Closest in time.
Eagle-3: Scaling up inference acceleration of large language models via training-time test (2025)
Li, Y., Wei, F., Zhang, C. & Zhang, H · 2025
Closest in time.
Transformers without normalization
Zhu, J., Chen, X., He, K., LeCun, Y. & Liu, Z · 2025
Closest in time.
The lottery llm hypothesis, rethinking what abilities should llm compression preserve?
Tang, Z. et al · 2025
Closest in time.