Fetching the paper…
Reading the bibliography…
Despite the remarkable capabilities, Large Language Models (LLMs) face deployment challenges due to their extensive size.
Genetic programming as a means for programming computers by natural selection
Koza, J. R · 1994
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W. J · 2015
Earlier work this paper cites.
Data-free parameter pruning for deep neural networks
Srinivas, S. and Babu, R. V · 2015
Earlier work this paper cites.
Extrapolation and learning equations
Martius, G. and Lampert, C. H · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Learning sparse neural networks through l_0 regularization
Louizos, C., Welling, M., and Kingma, D. P · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Efficient neural architecture search via parameter sharing
Pham, H., Guan, M. Y., Zoph, B., Le, Q. V., and Dean, J · 2018
Earlier work this paper cites.
Learning equations for extrapolation and control
Sahoo, S. S., Lampert, C. H., and Martius, G · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H · 2019
Earlier work this paper cites.
Rethinking the value of network pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Earlier work this paper cites.
Winogrande
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
The impact of gpu dvfs on the energy and performance of deep learning: An empirical study
Tang, Z., Wang, Y., Wang, Q., and Chu, X · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained bert networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Earlier work this paper cites.
Automl-zero: Evolving machine learning algorithms from scratch
Real, E., Liang, C., So, D., and Le, Q · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Earlier work this paper cites.
Green ai
Schwartz, R., Dodge, J., Smith, N. A., and Etzioni, O · 2020
Earlier work this paper cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H., Kunin, D., Yamins, D. L., and Ganguli, S · 2020
Earlier work this paper cites.
Communication-efficient distributed deep learning: A comprehensive survey
Tang, Z., Shi, S., Chu, X., Wang, W., and Li, B · 2020
Earlier work this paper cites.
Neural pruning via growing regularization
Wang, H., Qin, C., Zhang, Y., and Fu, Y. R · 2020
Earlier work this paper cites.
Pruning randomly initialized neural networks with iterative randomization
Chijiwa, D., Yamaguchi, S. y., Ida, Y., Umakoshi, K., and INOUE, T · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Cited alongside, same era.
Improving one-shot nas with shrinking-and-expanding supernet
Hu, Y., Wang, X., Li, L., and Gu, Q · 2021
Cited alongside, same era.
Werner, M., Junginger, A., Hennig, P., and Martius, G · 2021
Cited alongside, same era.
Red++ : Data-free pruning of deep neural networks via input splitting and output merging
Yvinec, E., Dapogny, A., Cord, M., and Bailly, K · 2021
Cited alongside, same era.
LLM-FP4: 4-bit floating-point quantized transformers
Liu, S.-y., Liu, Z., Huang, X., Dong, P., and Cheng, K.-T · 2023
Later among the works it cites.
Estimating the carbon footprint of bloom, a 176b parameter language model
Luccioni, A. S., Viguier, S., and Ligozat, A.-L · 2023
Later among the works it cites.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Later among the works it cites.
Lightllm
ModelTC · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Pruning pre-trained language models with principled importance and self-regularization
Ren, S. and Zhu, K · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
EZNAS: Evolving zero-cost proxies for neural architecture scoring
Akhauri, Y., Munoz, J. P., Jain, N., and Iyer, R · 2022
Cited alongside, same era.
Prior-guided one-shot neural architecture search
Dong, P., Niu, X., Li, L., Xie, L., Zou, W., Ye, T., Wei, Z., and Pan, H · 2022
Cited alongside, same era.
Nas-lid: Efficient neural architecture search with local intrinsic dimension
He, X., Yao, J., Wang, Y., Tang, Z., Cheung, K. C., See, S., Han, B., and Chu, X · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
A fast post-training pruning framework for transformers
Kwon, W., Kim, S., Mahoney, M. W., Hassoun, J., Keutzer, K., and Gholami, A · 2022
Cited alongside, same era.
Self-regulated feature learning via teacher-free feature distillation
Li, L · 2022
Cited alongside, same era.
Shadow knowledge distillation: Bridging offline and online knowledge transfer
Li, L. and Jin, Z · 2022
Cited alongside, same era.
Tang, Z., Wang, Y., He, X., Zhang, L., Pan, X., Wang, Q., Zeng, R., Zhao, K., Shi, S., He, B., et al · 2023
Later among the works it cites.
Tvt: Training-free vision transformer search on tiny datasets
Wei, Z., Pan, H., Li, L., Dong, P., Tian, Z., Niu, X., and Li, D · 2023
Later among the works it cites.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Later among the works it cites.
Norm: Knowledge distillation via n-to-one representation matching
Xiaolong, L., Lujun, L., Chao, L., and Yao, A · 2023
Later among the works it cites.
Perp: Rethinking the prune-retrain paradigm in the era of llms
Zimmer, M., Andoni, M., Spiegel, C., and Pokutta, S · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M., Jacobs, S. A., and etc., A. A. A · 2024
Closest in time.
SliceGPT: Compress large language models by deleting rows and columns
Ashkboos, S., Croci, M. L., do Nascimento, M. G., Hoefler, T., and Hensman, J · 2024
Closest in time.
Parzc: Parametric zero-cost proxies for efficient nas
Dong, P., Li, L., Pan, X., Wei, Z., Liu, X., Wang, Q., and Chu, X · 2024
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Closest in time.
Uniads: Universal architecture-distiller search for distillation gap
Lu, L., Chen, Z., Lu, X., Rao, Y., Li, L., and Pang, S · 2024
Closest in time.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
Sharma, P., Ash, J. T., and Misra, D · 2024
Closest in time.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2024
Closest in time.
Fedimpro: Measuring and improving client update in federated learning
Tang, Z., Zhang, Y., Shi, S., Tian, X., Liu, T., Han, B., and Chu, X · 2024
Closest in time.
The llm surgeon
van der Ouderaa, T. F. A., Nagel, M., van Baalen, M., Asano, Y. M., and Blankevoort, T · 2024
Closest in time.
Auto-prox: Training-free vision transformer architecture search via automatic proxy discovery
Wei, Z., Dong, P., Hui, Z., Li, A., Li, L., Lu, M., Pan, H., and Li, D · 2024
Closest in time.
BESA: Pruning large language models with blockwise parameter-efficient sparsity allocation
Xu, P., Shao, W., Chen, M., Tang, S., Zhang, K., Gao, P., An, F., Qiao, Y., and Luo, P · 2024
Closest in time.
Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning LLMs to high sparsity
Yin, L., Wu, Y., Zhang, Z., Hsieh, C.-Y., Wang, Y., Jia, Y., Pechenizkiy, M., Liang, Y., Wang, Z., and Liu, S · 2024
Closest in time.
Saswot: Real-time semantic segmentation architecture search without training
Zhu, C., Li, L., Wu, Y., and Sun, Z · 2024
Closest in time.
Learning transferable architectures for scalable image recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V · 2024
Closest in time.