Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have transformed the landscape of artificial intelligence, while their enormous size presents significant challenges in terms of computational costs.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., Chen, S., and Pan, S · 2017
Earlier work this paper cites.
Everitt, T., Lea, G., and Hutter, M · 2018
Earlier work this paper cites.
Openwebtext corpus, 2019
Aaron Gokaslan, Vanya Cohen, E. P. S. T · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Neural network compression via sparse optimization
Chen, T., Ji, B., Shi, Y., Ding, T., Fang, B., Yi, S., and Tu, X · 2020
Earlier work this paper cites.
Orthant based proximal stochastic gradient method for l1-regularized optimization
Chen, T., Ding, T., Ji, B., Wang, G., Shi, Y., Tian, J., Yi, S., Tu, X., and Zhu, Z · 2020
Earlier work this paper cites.
A standardized project gutenberg corpus for statistical analysis of natural language and quantitative linguistics
Gerlach, M. and Font-Clos, F · 2020
Earlier work this paper cites.
Cdfi: Compression-driven network design for frame interpolation
Ding, T., Liang, L., Zhu, Z., and Zharkov, I · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Sparsity-guided network design for frame interpolation
Ding, T., Liang, L., Zhu, Z., Chen, T., and Zharkov, I · 2022
Cited alongside, same era.
Parameter-efficient sparsity for large language models fine-tuning
A survey on large language models: Applications, challenges, limitations, and practical usage
Hadi, M. U., Qureshi, R., Shah, A., Irfan, M., Zafar, A., Shaikh, M. B., Akhtar, N., Wu, J., Mirjalili, S., et al · 2023
Closest in time.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Closest in time.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, Y., Luo, F., Tan, C., Wang, M., Huang, S., Li, S., and Bai, J · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Cited alongside, same era.
SparseGPT: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
Only train once: A one-shot neural network training and pruning framework
Chen, T., Ji, B., Ding, T., Fang, B., Wang, G., Zhu, Z., Liang, L., Shi, Y., Yi, S., and Tu, X
Cited in the paper.
Towards automatic neural architecture search within general super-networks
Chen, T., Liang, L., Ding, T., and Zharkov, I
Cited in the paper.
Otov2: Automatic, generic, user-friendly
Chen, T., Liang, L., Ding, T., Zhu, Z., and Zharkov, I
Cited in the paper.
Closest in time.
Sheared llama: Accelerating language model pre-training via structured pruning
Xia, M., Gao, T., Zeng, Z., and Chen, D · 2023
Closest in time.
Pruning meets low-rank parameter-efficient fine-tuning
Zhang, M., Shen, C., Yang, Z., Ou, L., Yu, X., Zhuang, B., et al · 2023
Closest in time.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Closest in time.