Fetching the paper…
Reading the bibliography…
A growing trend has emerged in designing high-quality Small Language Models (SLMs) with a few million parameters.
Quantizing for maximum output entropy (corresp.)
D. Messerschmitt · 1971
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
gemmlowp: A small self-contained low-precision gemm library
B. Jacob and P. Warden · 2017
Earlier work this paper cites.
Pact: Parameterized clipping activation for quantized neural networks
J. Choi, Z. Wang, S. Venkataramani, P. I.-J. Chuang, V. Srinivasan, and K. Gopalakrishnan · 2018
Earlier work this paper cites.
Qnnpack: Open source library for optimized mobile deep learning, 2018
M. Dukhan, Y. Wu, and H. Lu · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
BLiMP: The Benchmark of Linguistic Minimal Pairs for English
A. Warstadt, A. Parrish, H. Liu, A. Mohananey, W. Peng, S.-F. Wang, and S. R. Bowman · 2020
Earlier work this paper cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
H. Wang, Z. Zhang, and S. Han · 2021
Earlier work this paper cites.
Accelerating deep learning model inference on arm cpus with ultra-low bit quantization and runtime
S. Ashfaq, M. AskariHemmat, S. Sah, E. Saboori, O. Mastropietro, and A. Hoffman · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2022
Earlier work this paper cites.
Spvit: Enabling faster vision transformers via latency-aware soft token pruning
Z. Kong, P. Dong, X. Ma, X. Meng, W. Niu, M. Sun, X. Shen, G. Yuan, B. Ren, H. Tang, et al · 2022
Earlier work this paper cites.
Nipq: Noise injection pseudo quantization for automated dnn optimization
S. Park, J. So, J. Shin, and E. Park · 2022
Earlier work this paper cites.
Bibert: Accurate fully binarized bert
H. Qin, Y. Ding, M. Zhang, Q. Yan, A. Liu, Q. Dang, Z. Liu, and X. Liu · 2022
Cited alongside, same era.
Compression of generative pre-trained language models via quantization
C. Tao, L. Hou, W. Zhang, L. Shang, X. Jiang, Q. Liu, P. Luo, and N. Wong · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Cited alongside, same era.
Heatvit: Hardware-efficient adaptive token pruning for vision transformers
P. Dong, M. Sun, A. Lu, Y. Xie, K. Liu, Z. Kong, X. Meng, Z. Li, X. Lin, Z. Fang, et al · 2023
Cited alongside, same era.
Token fusion: Bridging the gap between token pruning and token merging
M. Kim, S. Gao, Y.-C. Hsu, Y. Shen, and H. Jin · 2024
Closest in time.
Token-scaled logit distillation for ternary weight generative language models
M. Kim, S. Lee, J. Lee, S. Hong, D.-S. Chang, W. Sung, and J. Choi · 2024
Closest in time.
Mobilellm: Optimizing sub-billion parameter language models for on-device use cases
Z. Liu, C. Zhao, F. Iandola, C. Lai, Y. Tian, I. Fedorov, Y. Xiong, E. Chang, Y. Shi, R. Krishnamoorthi, et al · 2024
Closest in time.
Agile-quant: Activation-guided quantization for faster inference of llms on the edge
X. Shen, P. Dong, L. Lu, Z. Kong, Z. Li, M. Lin, C. Wu, and Y. Wang · 2024
Closest in time.
Hotaq: Hardware oriented token adaptive quantization for large language models
X. Shen, Z. Han, L. Lu, Z. Kong, P. Dong, Z. Li, Y. Xie, C. Wu, M. Leeser, P. Zhao, X. Lin, and Y. Wang · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Kim, C. Hooper, A. Gholami, Z. Dong, X. Li, S. Shen, M. W. Mahoney, and K. Keutzer · 2023
Cited alongside, same era.
Awq: Activation-aware weight quantization for llm compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, and S. Han · 2023
Cited alongside, same era.
Llm-qat: Data-free quantization aware training for large language models
Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y. Mehdad, Y. Shi, R. Krishnamoorthi, and V. Chandra · 2023
Cited alongside, same era.
I. Timiryasov and J.-L. Tastet · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Cited alongside, same era.
Zeroquant-fp: A leap forward in llms post-training w4a8 quantization using floating-point formats
X. Wu, Z. Yao, and Y. He · 2023
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2023
Cited alongside, same era.
Search for efficient large language models
X. Shen, P. Zhao, Y. Gong, Z. Kong, Z. Zhan, Y. Wu, M. Lin, C. Wu, X. Lin, and Y. Wang · 2024
Closest in time.
Babyllama-2: Ensemble-distilled models consistently outperform teachers with limited data
J.-L. Tastet and I. Timiryasov · 2024
Closest in time.
A survey of small language models
C. Van Nguyen, X. Shen, R. Aponte, Y. Xia, S. Basu, Z. Hu, J. Chen, M. Parmar, S. Kunapuli, J. Barrow, et al · 2024
Closest in time.
Fully open source moxin-7b technical report
P. Zhao, X. Shen, Z. Kong, Y. Shen, S.-E. Chang, T. Rupprecht, L. Lu, E. Nan, C. Yang, Y. He, et al · 2024
Closest in time.
Smollm2: When smol goes big–data-centric training of a small language model
L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Blázquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydlíček, A. P. Lajarín, V. Srivastav, et al · 2025
Closest in time.
Draftattention: Fast video diffusion via low-resolution attention guidance
X. Shen, C. Han, Y. Zhou, Y. Xie, Y. Gong, Q. Wang, Y. Wang, Y. Wang, P. Zhao, and J. Gu · 2025
Closest in time.
Quartdepth: Post-training quantization for real-time depth estimation on the edge
X. Shen, W. Ma, J. Liu, C. Yang, R. Ding, Q. Wang, H. Ding, W. Niu, Y. Wang, P. Zhao, J. Lin, and J. Gu · 2025
Closest in time.
Fastcar: Cache attentive replay for fast auto-regressive video generation on the edge
X. Shen, W. Ma, Y. Zhou, E. Tang, Y. Xie, Z. Li, Y. Gong, Q. Wang, H. Ding, Y. Wang, Y. Wang, P. Zhao, J. Lin, and J. Gu · 2025
Closest in time.
Lazydit: Lazy learning for the acceleration of diffusion transformers
X. Shen, Z. Song, Y. Zhou, B. Chen, Y. Li, Y. Gong, K. Zhang, H. Tan, J. Kuen, H. Ding, Z. Shu, W. Niu, P. Zhao, Y. Wang, and J. Gu · 2025
Closest in time.
Numerical pruning for efficient autoregressive models
X. Shen, Z. Song, Y. Zhou, B. Chen, J. Liu, R. Zhang, R. A. Rossi, H. Tan, T. Yu, X. Chen, Y. Zhou, T. Sun, P. Zhao, Y. Wang, and J. Gu · 2025
Closest in time.
Sparse learning for state space models on mobile
X. Shen, H. Zheng, Y. Gong, Z. Kong, C. Yang, Z. Zhan, Y. Wu, X. Lin, Y. Wang, P. Zhao, and W. Niu · 2025
Closest in time.