Fetching the paper…
Reading the bibliography…
Transformer based large language models have achieved tremendous success.
Y. LeCun, J. Denker, and S. Solla, “Optimal brain damage,” Advances in neural information processing systems , vol. 2, 1989
1989
Earlier work this paper cites.
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural computation , vol. 3, no. 1, pp. 79–87, 1991
1991
Earlier work this paper cites.
B. Hassibi, D. G. Stork, and G. J. Wolff, “Optimal brain surgeon and general network pruning,” in IEEE international conference on neural networks . IEEE, 1993, pp. 293–299
1993
Earlier work this paper cites.
M. I. Jordan and R. A. Jacobs, “Hierarchical mixtures of experts and the em algorithm,” Neural computation , vol. 6, no. 2, pp. 181–214, 1994
1994
Earlier work this paper cites.
A. Rahimi and B. Recht, “Random features for large-scale kernel machines,” Advances in neural information processing systems , vol. 20, 2007
2007
Earlier work this paper cites.
A. Graves and A. Graves, “Long short-term memory,” Supervised sequence labelling with recurrent neural networks , pp. 37–45, 2012
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “Xnor-net: Imagenet classification using binary convolutional neural networks,” in European conference on computer vision . Springer, 2016, pp. 525–542
2016
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Yim, D. Joo, J.-H. Bae, and J. Kim, “A gift from knowledge distillation: Fast optimization, network minimization and transfer learning,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 7130–7138, 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:206596723
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 2704–2713
2018
Earlier work this paper cites.
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu, “Deep mutual learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 4320–4328
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu, “ERNIE: enhanced language representation with informative entities,” in Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , A. Korhonen, D. R. Traum, and L. Màrquez, Eds. Association for Computational Linguistics, 2019, pp. 1441–1451
2019
Earlier work this paper cites.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat, “Q8bert: Quantized 8bit bert,” in 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS) . IEEE, 2019, pp. 36–39
2019
Earlier work this paper cites.
K. Wang, Z. Liu, Y. Lin, J. Lin, and S. Han, “Haq: Hardware-aware automated quantization with mixed precision,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 8612–8620
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Elsken, J. H. Metzen, and F. Hutter, “Neural architecture search: A survey,” The Journal of Machine Learning Research , vol. 20, no. 1, pp. 1997–2017, 2019
2019
Earlier work this paper cites.
B. Wu, X. Dai, P. Zhang, Y. Wang, F. Sun, Y. Wu, Y. Tian, P. Vajda, Y. Jia, and K. Keutzer, “Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 10 734–10 742
2019
Earlier work this paper cites.
D. So, Q. Le, and C. Liang, “The evolved transformer,” in International conference on machine learning . PMLR, 2019, pp. 5877–5886
2019
Earlier work this paper cites.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International conference on machine learning . PMLR, 2019, pp. 6105–6114
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Molchanov, A. Mallya, S. Tyree, I. Frosio, and J. Kautz, “Importance estimation for neural network pruning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 11 264–11 272
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Q-bert: Hessian based ultra low precision quantization of bert,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 8815–8821
2020
Earlier work this paper cites.
W. Zhang, L. Hou, Y. Yin, L. Shang, X. Chen, X. Jiang, and Q. Liu, “Ternarybert: Distillation-aware ultra-low bit bert,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 509–521
2020
Earlier work this paper cites.
Z. Zhao, Y. Liu, L. Chen, Q. Liu, R. Ma, and K. Yu, “An investigation on different underlying quantization schemes for pre-trained language models,” in Natural Language Processing and Chinese Computing: 9th CCF International Conference, NLPCC 2020, Zhengzhou, China, October 14–18, 2020, Proceedings, Part I 9 . Springer, 2020, pp. 359–371
2020
Earlier work this paper cites.
A. H. Zadeh, I. Edo, O. M. Awad, and A. Moshovos, “Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 811–824
2020
Earlier work this paper cites.
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort, “Up or down? adaptive rounding for post-training quantization,” in International Conference on Machine Learning . PMLR, 2020, pp. 7197–7206
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Chen, J. Frankle, S. Chang, S. Liu, Y. Zhang, Z. Wang, and M. Carbin, “The lottery ticket hypothesis for pre-trained bert networks,” Advances in neural information processing systems , vol. 33, pp. 15 834–15 846, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
V. Sanh, T. Wolf, and A. Rush, “Movement pruning: Adaptive sparsity by fine-tuning,” Advances in Neural Information Processing Systems , vol. 33, pp. 20 378–20 389, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Zhang and Y. He, “Accelerating training of transformer-based language models with progressive layer dropping,” Advances in Neural Information Processing Systems , vol. 33, pp. 14 011–14 023, 2020
2020
Earlier work this paper cites.
S. Goyal, A. R. Choudhury, S. Raje, V. Chakaravarthy, Y. Sabharwal, and A. Verma, “Power-bert: Accelerating bert inference via progressive word-vector elimination,” in International Conference on Machine Learning . PMLR, 2020, pp. 3690–3699
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
L. Hou, Z. Huang, L. Shang, X. Jiang, X. Chen, and Q. Liu, “Dynabert: Dynamic bert with adaptive width and depth,” Advances in Neural Information Processing Systems , vol. 33, pp. 9782–9793, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, “Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers,” Advances in Neural Information Processing Systems , vol. 33, pp. 5776–5788, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. I. Mirzadeh, M. Farajtabar, A. Li, N. Levine, A. Matsukawa, and H. Ghasemzadeh, “Improved knowledge distillation via teacher assistant,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 04, 2020, pp. 5191–5198
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang et al. , “Big bird: Transformers for longer sequences,” Advances in neural information processing systems , vol. 33, pp. 17 283–17 297, 2020
2020
Earlier work this paper cites.
Y. Tay, D. Bahri, L. Yang, D. Metzler, and D.-C. Juan, “Sparse sinkhorn attention,” in International Conference on Machine Learning . PMLR, 2020, pp. 9438–9447
2020
Earlier work this paper cites.
X. Li, Y. Meng, M. Zhou, Q. Han, F. Wu, and J. Li, “Sac: Accelerating and structuring self-attention via sparse adaptive connection,” Advances in Neural Information Processing Systems , vol. 33, pp. 16 997–17 008, 2020
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International conference on machine learning . PMLR, 2020, pp. 5156–5165
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Wan, X. Dai, P. Zhang, Z. He, Y. Tian, S. Xie, B. Wu, M. Yu, T. Xu, K. Chen et al. , “Fbnetv2: Differentiable neural architecture search for spatial and channel dimensions,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 12 965–12 974
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Xin, R. Tang, J. Lee, Y. Yu, and J. J. Lin, “Deebert: Dynamic early exiting for accelerating bert inference,” in Annual Meeting of the Association for Computational Linguistics , 2020. [Online]. Available: https://api.semanticscholar.org/CorpusID:216552850
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020, pp. 3505–3506
2020
Earlier work this paper cites.
X. Jiang, H. Wang, Y. Chen, Z. Wu, L. Wang, B. Zou, Y. Yang, Z. Cui, Y. Cai, T. Yu et al. , “Mnn: A universal and efficient inference engine,” Proceedings of Machine Learning and Systems , vol. 2, pp. 1–13, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Lin, Y. Wang, X. Liu, and X. Qiu, “A survey of transformers,” CoRR , vol. abs/2106.04554, 2021
2021
Earlier work this paper cites.
Z. Zhang, X. Han, H. Zhou, P. Ke, Y. Gu, D. Ye, Y. Qin, Y. Su, H. Ji, J. Guan, F. Qi, X. Wang, Y. Zheng, G. Zeng, H. Cao, S. Chen, D. Li, Z. Sun, Z. Liu, M. Huang, W. Han, J. Tang, J. Li, X. Zhu, and M. Sun, “CPM: A large-scale generative chinese pre-trained language model,” AI Open , vol. 2, pp. 93–99, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Qin, Y. Ding, M. Zhang, Y. Qinghua, A. Liu, Q. Dang, Z. Liu, and X. Liu, “Bibert: Accurate fully binarized bert,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
B. Wang, Y. Ren, L. Shang, X. Jiang, and Q. Liu, “Exploring extreme parameter compression for pre-trained language models,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
H. Bai, W. Zhang, L. Hou, L. Shang, J. Jin, X. Jiang, Q. Liu, M. Lyu, and I. King, “Binarybert: Pushing the limit of bert quantization,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 4334–4348
2021
Earlier work this paper cites.
S. Kim, A. Gholami, Z. Yao, M. W. Mahoney, and K. Keutzer, “I-bert: Integer-only bert quantization,” in International conference on machine learning . PMLR, 2021, pp. 5506–5518
2021
Earlier work this paper cites.
S. Dai, R. Venkatesan, M. Ren, B. Zimmer, W. Dally, and B. Khailany, “Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,” Proceedings of Machine Learning and Systems , vol. 3, pp. 873–884, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. Dettmers, M. Lewis, S. Shleifer, and L. Zettlemoyer, “8-bit optimizers via block-wise quantization,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
B. Cui, Y. Li, and Z. Zhang, “Joint structured pruning and dense knowledge distillation for efficient transformer model compression,” Neurocomputing , vol. 458, pp. 56–69, 2021
2021
Earlier work this paper cites.
J. Li, R. Cotterell, and M. Sachan, “Differentiable subset pruning of transformer heads,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1442–1459, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Liu, F. Li, G. Li, and J. Cheng, “Ebert: Efficient bert inference with dynamic structured pruning,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 4814–4823
2021
Earlier work this paper cites.
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2021, pp. 97–110
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Ren, H. Dai, Z. Dai, M. Yang, J. Leskovec, D. Schuurmans, and B. Dai, “Combiner: Full attention transformer with sparse computation cost,” Advances in Neural Information Processing Systems , vol. 34, pp. 22 470–22 482, 2021
2021
Earlier work this paper cites.
A. Roy, M. Saffar, A. Vaswani, and D. Grangier, “Efficient content-based sparse attention with routing transformers,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 53–68, 2021
2021
Earlier work this paper cites.
Y. Xiong, Z. Zeng, R. Chakraborty, M. Tan, G. Fung, Y. Li, and V. Singh, “Nyströmformer: A nyström-based algorithm for approximating self-attention,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 16, 2021, pp. 14 138–14 148
2021
Earlier work this paper cites.
Z. Shen, M. Zhang, H. Zhao, S. Yi, and H. Li, “Efficient attention: Attention with linear complexities,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2021, pp. 3531–3539
2021
Earlier work this paper cites.
P. Ren, Y. Xiao, X. Chang, P.-Y. Huang, Z. Li, X. Chen, and X. Wang, “A comprehensive survey of neural architecture search: Challenges and solutions,” ACM Computing Surveys (CSUR) , vol. 54, no. 4, pp. 1–34, 2021
2021
Cited alongside, same era.
Y. Liu, Y. Sun, B. Xue, M. Zhang, G. G. Yen, and K. C. Tan, “A survey on evolutionary neural architecture search,” IEEE transactions on neural networks and learning systems , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2023
Later among the works it cites.
T. Pegolotti, E. Frantar, D. Alistarh, and M. Püschel, “Generating efficient kernels for quantized inference on large language models,” in Workshop on Efficient Systems for Foundation Models@ ICML2023 , 2023
2023
Later among the works it cites.
C. Yu, T. Chen, and Z. Gan, “Boost transformer-based language models with gpu-friendly sparsity and quantization,” in Findings of the Association for Computational Linguistics: ACL 2023 , 2023, pp. 218–235
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Zhu, “Leebert: Learned early exit for bert with cross-level optimization,” in Annual Meeting of the Association for Computational Linguistics , 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:236459809
2021
Cited alongside, same era.
D. Ye, Y. Lin, Y. Huang, and M. Sun, “Tr-bert: Dynamic token reduction for accelerating bert inference,” in North American Chapter of the Association for Computational Linguistics , 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:235097557
2021
Cited alongside, same era.
2021
Cited alongside, same era.
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer, “Base layers: Simplifying training of large, sparse models,” in International Conference on Machine Learning . PMLR, 2021, pp. 6265–6274
2021
Cited alongside, same era.
S. Roller, S. Sukhbaatar, J. Weston et al. , “Hash layers for large sparse models,” Advances in Neural Information Processing Systems , vol. 34, pp. 17 555–17 566, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby, “Scaling vision with sparse mixture of experts,” Advances in Neural Information Processing Systems , vol. 34, pp. 8583–8595, 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
B. Isik, H. Kumbong, W. Ning, X. Yao, S. Koyejo, and C. Zhang, “Gpt-zip: Deep compression of finetuned large language models,” in Workshop on Efficient Systems for Foundation Models@ ICML2023 , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
T. Dettmers and L. Zettlemoyer, “The case for 4-bit precision: k-bit inference scaling laws,” in International Conference on Machine Learning . PMLR, 2023, pp. 7750–7774
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. K. Jaiswal, S. Liu, T. Chen, Y. Ding, and Z. Wang, “Instant soup: Cheap pruning ensembles in a single pass can draw lottery tickets from large models,” in International Conference on Machine Learning . PMLR, 2023, pp. 14 691–14 701
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Klein, J. Golebiowski, X. Ma, V. Perrone, and C. Archambeau, “Structural pruning of large language models via neural architecture search,” in AutoML Conference 2023 (Workshop) , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Yang, Y. Jang, H. Lee, S. Jeong, and K. Jung, “Task-specific compression for multi-task language models using attribution-based pruning,” in Findings of the Association for Computational Linguistics: EACL 2023 , 2023, pp. 582–592
2023
Later among the works it cites.
C. Tao, L. Hou, H. Bai, J. Wei, X. Jiang, Q. Liu, P. Luo, and N. Wong, “Structured pruning for efficient generative pre-trained language models,” in Findings of the Association for Computational Linguistics: ACL 2023 , 2023, pp. 10 880–10 895
2023
Later among the works it cites.
H. Sajjad, F. Dalvi, N. Durrani, and P. Nakov, “On the effect of dropping layers of pre-trained transformer models,” Computer Speech & Language , vol. 77, p. 101429, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Zhang, H. Bai, H. Lin, J. Zhao, L. Hou, and C. V. Cannistraci, “An efficient plug-and-play post-training pruning strategy in large language models,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
E. Frantar and D. Alistarh, “Sparsegpt: Massive language models can be accurately pruned in one-shot,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Anonymous, “Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning LLMs to high sparsity,” in Submitted to The Twelfth International Conference on Learning Representations , 2023, under review. [Online]. Available: https://openreview.net/forum?id=pOBvr1PxFd
2023
Later among the works it cites.
A. Syed, P. H. Guo, and V. Sundarapandiyan, “Prune and tune: Improving efficient pruning techniques for massive language models,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Zhang, A. Muhamed, A. Anantharaman, G. Wang, C. Chen, K. Zhong, Q. Cui, Y. Xu, B. Zeng, T. M. Chilimbi, and Y. Chen, “Reaugkd: Retrieval-augmented knowledge distillation for pre-trained language models,” in Annual Meeting of the Association for Computational Linguistics , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:259370551
2023
Later among the works it cites.
2023
Later among the works it cites.
C. Liang, S. Zuo, Q. Zhang, P. He, W. Chen, and T. Zhao, “Less is more: Task-aware layer-wise distillation for language model compression,” in International Conference on Machine Learning . PMLR, 2023, pp. 20 852–20 867
2023
Later among the works it cites.
S. Dasgupta, T. Cohn, and T. Baldwin, “Cost-effective distillation of large language models,” in Annual Meeting of the Association for Computational Linguistics , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:259858962
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford_alpaca , 2023
2023
Later among the works it cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
Later among the works it cites.
Y. Anand, Z. Nussbaum, B. Duderstadt, B. Schmidt, and A. Mulyar, “Gpt4all: Training an assistant-style chatbot with large scale data distillation from gpt-3.5-turbo,” GitHub , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Shridhar, A. Stolfo, and M. Sachan, “Distilling reasoning capabilities into smaller language models,” in Findings of the Association for Computational Linguistics: ACL 2023 , 2023, pp. 7059–7073
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Chen, Q. Gao, A. Bosselut, A. Sabharwal, and K. Richardson, “Disco: distilling counterfactuals with large language models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 5514–5528
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Wang, K. Chen, H. Tan, and K. Guo, “Tabi: An efficient multi-level inference system for large language models,” Proceedings of the Eighteenth European Conference on Computer Systems , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:258508784
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Patel and G. Wong, “Gpt-4 architecture, infrastructure, training dataset, costs, vision, moe,” 2023, https://www.semianalysis.com/p/gpt-4-architecture-infrastructure
2023
Later among the works it cites.
Y. Zhou, N. Du, Y. Huang, D. Peng, C. Lan, D. Huang, S. Shakeri, D. So, A. M. Dai, Y. Lu et al. , “Brainformers: Trading simplicity for efficiency,” in International Conference on Machine Learning . PMLR, 2023, pp. 42 531–42 542
2023
Later among the works it cites.
Y. Xie, S. Huang, T. Chen, and F. Wei, “Moec: Mixture of expert clusters,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 11, 2023, pp. 13 807–13 815
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Diao, T. Xu, R. Xu, J. Wang, and T. Zhang, “Mixture-of-domain-adapters: Decoupling and injecting domain knowledge to pre-trained language models’ memories,” in Annual Meeting of the Association for Computational Linguistics , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:259108831
2023
Later among the works it cites.
R. Li, G. Murray, and G. Carenini, “Mixture-of-linguistic-experts adapters for improving and interpreting pre-trained language models,” in Conference on Empirical Methods in Natural Language Processing , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:264487239
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Zhai, C. Jiang, L. Wang, X. Jia, S. Zhang, Z. Chen, X. Liu, and Y. Zhu, “Bytetransformer: A high-performance transformer boosted for variable-length inputs,” in 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2023, pp. 344–355
2023
Later among the works it cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single gpu,” in International Conference on Machine Learning . PMLR, 2023, pp. 31 094–31 116
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Li, H. Liu, Z. Bian, J. Fang, H. Huang, Y. Liu, B. Wang, and Y. You, “Colossal-ai: A unified deep learning system for large-scale parallel training,” in Proceedings of the 52nd International Conference on Parallel Processing , 2023, pp. 766–775
2023
Later among the works it cites.
A. Pham, C. Yang, S. Sheng, S. Zhao, S. Lee, B. Jiang, F. Dong, X. Guan, and F. Ming, “Openllm: Operating llms in production,” 2023, https://github.com/bentoml/OpenLLM
2023
Later among the works it cites.
M. team, “MLC-LLM,” 2023. [Online]. Available: https://github.com/mlc-ai/mlc-llm
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Anonymous, “Pushing gradient towards zero: A novel pruning method for large language models,” 2024. [Online]. Available: https://openreview.net/forum?id=IU4L7wiwxw
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
——, “BESA: Pruning large language models with blockwise parameter-efficient sparsity allocation,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=gC6JTEU3jl
2024
Closest in time.
V. Boža, “Fast and optimal weight update for pruned large language models,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Li, Z. Sun, X. He, L. Zeng, Y. Lin, E. Li, B. Zheng, R. Zhao, and X. Chen, “Locmoe: A low-overhead moe for large language model training,” 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:267212059
2024
Closest in time.
F. Xue, Z. Zheng, Y. Fu, J. Ni, Z. Zheng, W. Zhou, and Y. You, “Openmoe: An early effort on open mixture-of-experts language models,” 2024
2024
Closest in time.