Fetching the paper…
Reading the bibliography…
As foundation AI models continue to increase in size, an important question arises - is massive scale the only path forward? This survey of about 160 papers presents a family of Small Language Models (SLMs) in the 1 to 8 billion parameter range that demonstrate smaller models can perform as well, or even outperform large models.
Distilling the knowledge in a neural network
Hinton, G., et al · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., et al · 2017
Earlier work this paper cites.
Can a suit of armor conduct electricity? A new dataset for open book question answering
Mihaylov, T., et al · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W.W., and Lu, X · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., et al · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
AdapterHub: A framework for adapting transformers
Pfeiffer, J., Rücklé, A., Poth, C., Kamath, A., Vulić, I., Ruder, S., Cho, K., and Gurevych, I · 2020
Earlier work this paper cites.
WinoGrande: An adversarial Winograd schema challenge at scale
Sakaguchi, K., Le Bras, R., Bhagavatula, C., and Choi, Y · 2020
Earlier work this paper cites.
Mish: A self regularized non-monotonic activation function
Misra, D · 2020
Earlier work this paper cites.
PIQA: Reasoning about physical commonsense in natural language
Bisk, Y., et al · 2020
Earlier work this paper cites.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
Kurtz, M., et al · 2020
Earlier work this paper cites.
FinBERT: A pre-trained financial language representation model for financial text mining
Liu, Z., et al · 2020
Earlier work this paper cites.
XtremeDistil: Multi-stage distillation for massive multilingual models
Mukherjee, S. and Awadallah, A · 2020
Earlier work this paper cites.
MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X.L. and Liang, P · 2021
Earlier work this paper cites.
AdapterFusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2021
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., et al · 2021
Earlier work this paper cites.
Compacter: Efficient low-rank hypercomplex adapter layers
Karimi Mahabadi, R., Henderson, J., and Ruder, S · 2021
Earlier work this paper cites.
Learn-to-share: A hardware-friendly transfer learning framework exploiting computation and parameter sharing
Fu, C., et al · 2021
Earlier work this paper cites.
Knowledge distillation: A survey
Gou, J., et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., et al · 2021
Earlier work this paper cites.
It’s not just size that matters: Small language models are also few-shot learners
Schick, T. and Schütze, H · 2021
Earlier work this paper cites.
DatlMedQA: A data augmentation and transfer learning based solution for medical question answering
Zhou, S. and Zhang, Y · 2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2022
Earlier work this paper cites.
Advancing model pruning via bi-level optimization
Zhang, Y., Yao, Y., Ram, P., Zhao, P., Chen, T., Hong, M., Wang, Y., and Liu, S · 2022
Earlier work this paper cites.
UniPELT: A unified framework for parameter-efficient language model tuning
Mao, Y., Mathias, L., Hou, R., Almahairi, A., Ma, H., Han, J., Yih, W., and Khabsa, M · 2022
Earlier work this paper cites.
AdaMix: Mixture-of-adapters for parameter-efficient tuning of large language models
Wang, Y., Mukherjee, S., Liu, X., Gao, J., Awadallah, A.H., and Gao, J · 2022
Earlier work this paper cites.
Data augmentation for biomedical factoid question answering
Pappas, D., Malakasiotis, P., and Androutsopoulos, I · 2022
Earlier work this paper cites.
Attention Fusion: a light yet efficient late fusion mechanism for task adaptation in NLU
Cao, J., Satya Prakash, C., and Hamza, W · 2022
Earlier work this paper cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., et al · 2022
Earlier work this paper cites.
Improved knowledge distillation for pre-trained language models via knowledge selection
Wang, C., et al · 2022
Earlier work this paper cites.
CodeGen: An open large language model for code with multi-turn program synthesis
Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C · 2023
Earlier work this paper cites.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J.W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Earlier work this paper cites.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C.D., and Finn, C · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Earlier work this paper cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J.E., and Stoica, I · 2023
Earlier work this paper cites.
Krona: Parameter efficient tuning with Kronecker adapter
Edalati, A., Tahaei, M., Kobyzev, I., Nia, V.P., Clark, J.J., and Rezagholizadeh, M · 2023
Earlier work this paper cites.
Teaching small language models to reason
Magister, L.C., Mallinson, J., Adamek, J., Malmi, E., and Severyn, A · 2023
Earlier work this paper cites.
Accelerating transformer inference for translation via parallel decoding
Santilli, A., Severino, S., Postolache, E., Maiorca, V., Mancusi, M., Marin, R., and Rodolà, E · 2023
Cited alongside, same era.
RWKV: Reinventing RNNs for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K.K., et al · 2023
Cited alongside, same era.
CodeGeeX: A pre-trained model for code generation with multilingual benchmarking on HumanEval-X
Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Shen, L., Wang, Z., Wang, A., Li, Y., Su, T., Yang, Z., and Tang, J · 2023
Cited alongside, same era.
PASS: Parallel speculative sampling
Monea, G., Joulin, A., and Grave, E · 2023
Cited alongside, same era.
Accelerating LLM inference with staged speculative decoding
Spector, B. and Re, C · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
OpenELM: An efficient language model family with open training and inference framework
Mehta, S., Sekhavat, M.H., Cao, Q., Horton, M., Jin, Y., Sun, C., Mirzadeh, I., Najibi, M., Belenko, D., Zatloukal, P., and Rastegari, M · 2024
Later among the works it cites.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws
Sardana, N. and Frankle, J · 2024
Later among the works it cites.
Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation
Xu, H., Sharaf, A., Chen, Y., Tan, W., Shen, L., Van Durme, B., Murray, K., and Kim, Y.J · 2024
Later among the works it cites.
AWQ: Activation-aware weight quantization for LLM compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., Gan, C., and Han, S · 2024
Later among the works it cites.
Zephyr: Direct distillation of LM alignment
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A.M., and Wolf, T · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Srivastava, A., et al · 2023
Cited alongside, same era.
StarCoder: May the source be with you!
Li, R., Allal, L.B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., et al · 2023
Cited alongside, same era.
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C.C.T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P.C., de Rosa, G.H., Saarikivi, O., et al · 2023
Cited alongside, same era.
Orca 2: Teaching small language models how to reason
Mitra, A., Del Corro, L., Mahajan, S., Codas, A., Simoes, C., Agarwal, S., Chen, X., Razdaibiedina, A., Jones, E., Aggarwal, K., et al · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., et al · 2023
Cited alongside, same era.
Code Llama: Open foundation models for code
Rozière, B., et al · 2023
Cited alongside, same era.
Parameter-efficient fine-tuning design spaces
Chen, J., et al · 2023
Cited alongside, same era.
TinyAgent: Function calling at the edge
Erdogan, L.E., Lee, N., Jha, S., Kim, S., Tabrizi, R., Moon, S., Hooper, C.R.C., Anumanchipalli, G., Keutzer, K., and Gholami, A · 2024
Later among the works it cites.
Granite-Function Calling Model: Introducing function calling abilities via multi-task learning of granular tasks
Abdelaziz, I., Basu, K., Agarwal, M., Kumaravel, S., Stallone, M., Panda, R., Rizk, Y., Bhargav, G.P.S., Crouse, M., Gunasekara, C., et al · 2024
Later among the works it cites.
Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding
Xia, H., Yang, Z., Dong, Q., Wang, P., Li, Y., Ge, T., Liu, T., Li, W., and Sui, Z · 2024
Later among the works it cites.
AGIEval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2024
Later among the works it cites.
Break the sequential dependency of LLM inference using lookahead decoding
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2024
Later among the works it cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Ponde, H. P. d. O. Pinto, Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., et al · 2024
Later among the works it cites.
RoFormer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2024
Later among the works it cites.
A comprehensive survey on pretrained foundation models: A history from BERT to ChatGPT
Zhou, C., Li, Q., Li, C., Yu, J., Liu, Y., Wang, G., Zhang, K., Ji, C., Yan, Q., He, L., et al · 2024
Later among the works it cites.
Dubey, A., et al · 2024
Later among the works it cites.
Zamba: A compact 7B SSM hybrid model
Glorioso, P., et al · 2024
Later among the works it cites.
SLADE: A portable small language model decompiler for optimized assembly
Armengol-Estapé, J., Woodruff, J., Cummins, C., and O’Boyle, M.F.P · 2024
Later among the works it cites.
Speculative streaming: Fast LLM inference without auxiliary models
Bhendawade, N., et al · 2024
Later among the works it cites.
Medusa: Simple LLM inference acceleration framework with multiple decoding heads
Cai, T., et al · 2024
Later among the works it cites.
SpeedUpNet: A plug-and-play hyper-network for accelerating text-to-image diffusion models
Chai, W., et al · 2024
Later among the works it cites.
ULTRAFEEDBACK: Boosting language models with scaled AI feedback
Cui, G., Yuan, L., Ding, N., Yao, G., He, B., Zhu, W., Ni, Y., Xie, G., Xie, R., Lin, Y., Liu, Z., and Sun, M · 2024
Later among the works it cites.
PowerInfer: Fast large language model serving with a consumer-grade GPU
Song, Y., et al · 2024
Later among the works it cites.
A survey on large language models: Applications, challenges, limitations, and practical usage
Hadi, M.U., et al · 2024
Later among the works it cites.
Neural machine translation of clinical text: An empirical investigation into multilingual pre-trained language models and transfer-learning
Han, L., et al · 2024
Later among the works it cites.
FLAME: A small language model for spreadsheet formulas
Joshi, H., et al · 2024
Later among the works it cites.
Can small language models help large language models reason better?: LM-guided chain-of-thought
Lee, J., et al · 2024
Later among the works it cites.
Prometheus-vision: Vision-language model as a judge for fine-grained evaluation
Lee, S., et al · 2024
Later among the works it cites.
DoRA: Weight-decomposed low-rank adaptation
Liu, S.-Y., et al · 2024
Later among the works it cites.
WizardCoder: Empowering code large language models with Evol-Instruct
Luo, Z., et al · 2024
Later among the works it cites.
Closer look at efficient inference methods: A survey of speculative decoding
Ryu, H. and Kim, E · 2024
Later among the works it cites.
MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., et al · 2024
Later among the works it cites.
OLMo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A., Ivison, H., Magnusson, I., Wang, Y., et al · 2024
Later among the works it cites.
WizardMath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D · 2025
Closest in time.
Octopus: On-device language model for function calling of software APIs
Chen, W., Li, Z., and Ma, M · 2025
Closest in time.
Efficient LLM inference on CPUs
Shen, H., Chang, H., Dong, B., Luo, Y., and Meng, H · 2025
Closest in time.
ToolACE: Winning the points of LLM function calling
Liu, W., Huang, X., Zeng, X., Hao, X., Yu, S., Li, D., Wang, S., Gan, W., Liu, Z., Yu, Y., et al · 2025
Closest in time.
Small language models learn enhanced reasoning skills from medical textbooks
Kim, H., Hwang, H., Lee, J., Park, S., Kim, D., Lee, T., Yoon, C., Sohn, J., Choi, D., Chen, Q., et al · 2025
Closest in time.
Jamba: Hybrid transformer-mamba language models
Lieber, O., et al · 2025
Closest in time.
xLAM: A family of large action models to empower AI agent systems
Zhang, J., et al · 2025
Closest in time.
Hymba: A hybrid-head architecture for small language models
Dong, X., et al · 2025
Closest in time.
A comprehensive overview of large language models
Naveed, H., Khan, A.U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., and Mian, A · 2025
Closest in time.
VideoChat: Chat-centric video understanding
Li, K.Q., et al · 2025
Closest in time.
SmolLM2: When Smol Goes Big – Data-centric training of a small language model
Ben Allal, L., et al · 2025
Closest in time.
Yang, A., et al · 2025
Closest in time.
GSM-Symbolic: Understanding the limitations of mathematical reasoning in large language models
Mirzadeh, S.I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M · 2025
Closest in time.