Fetching the paper…
Reading the bibliography…
The interest in developing small language models (SLM) for on-device deployment is fast growing.
Decoupled weight decay regularization
Loshchilov, I · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Le Bras, R., Bhagavatula, C., and Choi, Y · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
bigscience/bloom-560m
BigScience · 2022
Earlier work this paper cites.
bigscience/bloomz-1b1
BigScience · 2022
Earlier work this paper cites.
facebook/opt-1.3b
Facebook · 2022
Earlier work this paper cites.
facebook/opt-125m
Facebook · 2022
Earlier work this paper cites.
facebook/galactica-1.3b
Facebook · 2022
Earlier work this paper cites.
The stack: 3 tb of permissively licensed source code
Kocetkov, D., Li, R., Allal, L. B., Li, J., Mou, C., Ferrandis, C. M., Jernite, Y., Mitchell, M., Hughes, S., Wolf, T., et al · 2022
Earlier work this paper cites.
Llm in a flash: Efficient large language model inference with limited memory
Alizadeh, K., Mirzadeh, I., Belenko, D., Khatamifard, K., Cho, M., Del Mundo, C. C., Rastegari, M., and Farajtabar, M · 2023
Earlier work this paper cites.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Earlier work this paper cites.
cerebras/cerebras-gpt-1.3b
Cerebras · 2023
Cited alongside, same era.
cerebras/cerebras-gpt-590m
Cerebras · 2023
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B · 2023
Cited alongside, same era.
Eleutherai/pythia-1.4b
EleutherAI · 2023
Cited alongside, same era.
Eleutherai/pythia-410m
EleutherAI · 2023
Cited alongside, same era.
Starcoder: may the source be with you!
Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., et al · 2023
Cited alongside, same era.
Mammoth: Building math generalist models through hybrid instruction tuning
Yue, X., Qu, X., Zhang, G., Fu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Later among the works it cites.
Qwen 1.5
Alibaba · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Language model evaluation harness, 2024
EleutherAI · 2024
Closest in time.
glaive function calling, 2024
glaiveai · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification, 2023
Lian, W., Wang, G., Goodson, B., Pentland, E., Cook, A., Vong, C., and ”Teknium” · 2023
Cited alongside, same era.
Deja vu: Contextual sparsity for efficient llms at inference time
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., et al · 2023
Cited alongside, same era.
Mobilellama
Meituan · 2023
Cited alongside, same era.
Octopack: Instruction tuning code large language models
Muennighoff, N., Liu, Q., Zebaze, A., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., von Werra, L., and Longpre, S · 2023
Cited alongside, same era.
Openwebmath: An open dataset of high-quality mathematical web text, 2023
Paster, K., Santos, M. D., Azerbayev, Z., and Ba, J · 2023
Cited alongside, same era.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Song, Y., Mi, Z., Xie, H., and Chen, H · 2023
Cited alongside, same era.
Hu, S., Tu, Y., Han, X., He, C., Cui, G., Long, X., Zheng, Z., Fang, Y., Huang, Y., Zhao, W., et al · 2024
Closest in time.
Magicoder evol instruct, 2024
ise uiuc · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., et al · 2024
Closest in time.
Datacomp-lm: In search of the next generation of training sets for language models, 2024
Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S., Bansal, H., Guha, E., Keh, S., Arora, K., Garg, S., Xin, R., Muennighoff, N., Heckel, R., Mercat, J., Chen, M., Gururangan, S., Wortsman, M., Albalak, A., Bitton, Y., Nezhurina, M., Abbas, A., Hsieh, C.-Y., Ghosh, D., Gardner, J., Kilian, M., Zhang, H., Shao, R., Pratt, S., Sanyal, S., Ilharco, G., Daras, G., Marathe, K., Gokaslan, A., Zhang, J., Chandu, K., Nguyen, T., Vasiljevic, I., Kakade, S., Song, S., Sanghavi, S., Faghri, F., Oh, S., Zettlemoyer, L., Lo, K., El-Nouby, A., Pouransari, H., Toshev, A., Wang, S., Groeneveld, D., Soldaini, L., Koh, P. W., Jitsev, J., Kollar, T., Dimakis, A. G., Carmon, Y., Dave, A., Schmidt, L., and Shankar, V · 2024
Closest in time.
Apigen: Automated pipeline for generating verifiable and diverse function-calling datasets
Liu, Z., Hoang, T., Zhang, J., Zhu, M., Lan, T., Kokane, S., Tan, J., Yao, W., Liu, Z., Feng, Y., et al · 2024
Closest in time.
Small language models: Survey, measurements, and insights
Lu, Z., Li, X., Cai, D., Yi, R., Liu, F., Zhang, X., Lane, N. D., and Xu, M · 2024
Closest in time.
Mobillama
MBZUAI · 2024
Closest in time.
microsoft/phi-3-mini
Microsoft · 2024
Closest in time.
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Lambert, N., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, P., Zettlemoyer, L., Smith, N. A., Hajishirzi, H., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K · 2024
Closest in time.
stabilityai/stablelm-2-zephyr*
StabilityAI · 2024
Closest in time.
Llm as a system service on mobile devices
Yin, W., Xu, M., Li, Y., and Liu, X · 2024
Closest in time.
Mobile foundation model as firmware
Yuan, J., Yang, C., Cai, D., Wang, S., Yuan, X., Zhang, Z., Li, X., Zhang, D., Mei, H., Jia, X., et al · 2024
Closest in time.