Fetching the paper…
Reading the bibliography…
Modern large language models (LLMs) driven by scaling laws, achieve intelligence emergency in large model sizes.
Constrained differential optimization
Platt, J. and Barr, A · 1987
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Hassibi, B., Stork, D. G., and Wolff, G. J · 1993
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Compresso: Pragmatic main memory compression
Choukse, E., Erez, M., and Alameldeen, A. R · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Darts: Differentiable architecture search
Liu, H., Simonyan, K., and Yang, Y · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Learning transferable architectures for scalable image recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V · 2018
Earlier work this paper cites.
Once-for-all: Train one network and specialize it for efficient deployment
Cai, H., Gan, C., Wang, T., Zhang, Z., and Han, S · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search
Wu, B., Dai, X., Zhang, P., Wang, Y., Sun, F., Wu, Y., Tian, Y., Vajda, P., Jia, Y., and Keutzer, K · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Bignas: Scaling up neural architecture search with big single-stage models
Yu, J., Jin, P., Liu, H., Bender, G., Kindermans, P.-J., Tan, M., Huang, T., Song, X., Pang, R., and Le, Q · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Earlier work this paper cites.
Nas-bert: Task-agnostic and adaptive-size bert compression with neural architecture search
Xu, J., Tan, X., Luo, R., Song, K., Li, J., Qin, T., and Liu, T.-Y · 2021
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al · 2022
Cited alongside, same era.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Frantar, E. and Alistarh, D · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Smollm - blazingly fast and remarkably powerful, 2024
Allal, L. B., Lozhkov, A., Bakouch, E., von Werra, L., and Wolf, T · 2024
Later among the works it cites.
Finding transformer circuits with edge pruning
Bhaskar, A., Wettig, A., Friedman, D., and Chen, D · 2024
Later among the works it cites.
Pruner-zero: Evolving symbolic pruning metric from scratch for large language models
Dong, P., Li, L., Tang, Z., Liu, X., Pan, X., Wang, Q., and Chu, X · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
A framework for few-shot language model evaluation, 07 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Cited alongside, same era.
Lorashear: Efficient large language model structured pruning and knowledge recovery
Chen, T., Ding, T., Yadav, B., Zharkov, I., and Liang, L · 2023
Cited alongside, same era.
Lighteval: A lightweight framework for llm evaluation, 2023
Fourrier, C., Habib, N., Wolf, T., and Tunstall, L · 2023
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
Le Scao, T., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models, 2023
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Later among the works it cites.
Olmo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., et al · 2024
Later among the works it cites.
Minillm: Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2024
Later among the works it cites.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
Later among the works it cites.
Distillm: Towards streamlined distillation for large language models
Ko, J., Kim, S., Chen, T., and Yun, S.-Y · 2024
Later among the works it cites.
Datacomp-lm: In search of the next generation of training sets for language models, 2024
Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S., Bansal, H., Guha, E., Keh, S., Arora, K., et al · 2024
Later among the works it cites.
Nuteprune: Efficient progressive pruning with numerous teachers for large language models
Li, S., Chen, J., Han, X., and Bai, J · 2024
Later among the works it cites.
Mobilellm: Optimizing sub-billion parameter language models for on-device use cases
Liu, Z., Zhao, C., Iandola, F., Lai, C., Tian, Y., Fedorov, I., Xiong, Y., Chang, E., Shi, Y., Krishnamoorthi, R., et al · 2024
Later among the works it cites.
Fineweb-edu: the finest collection of educational content, 2024
Lozhkov, A., Ben Allal, L., von Werra, L., and Wolf, T · 2024
Later among the works it cites.
Kvpruner: Structural pruning for faster and memory-efficient large language models
Lv, B., Zhou, Q., Ding, X., Wang, Y., and Ma, Z · 2024
Later among the works it cites.
Llm pruning and distillation in practice: The minitron approach
Sreenivas, S. T., Muralidharan, S., Joshi, R., Chochowski, M., Patwary, M., Shoeybi, M., Catanzaro, B., Kautz, J., and Molchanov, P · 2024
Later among the works it cites.
Rethinking optimization and architecture for tiny language models
Tang, Y., Liu, F., Ni, Y., Tian, Y., Bai, Z., Hu, Y.-Q., Liu, S., Jui, S., Han, K., and Wang, Y · 2024
Later among the works it cites.
Redpajama: an open dataset for training large language models
Weber, M., Fu, D., Anthony, Q., Oren, Y., Adams, S., Alexandrov, A., Lyu, X., Nguyen, H., Yao, X., Adams, V., et al · 2024
Later among the works it cites.
Apt: Adaptive pruning and tuning pretrained language models for efficient training and inference
Zhao, B., Hajishirzi, H., and Cao, Q · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.