Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have attracted extensive attention due to their remarkable performance across various tasks.
J. Xu, X. Tan, R. Luo, K. Song, J. Li, T. Qin, and T.-Y. Liu, “Nas-bert: task-agnostic and adaptive-size bert compression with neural architecture search,” in Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining , 2021, pp. 1933–1943
1943
Earlier work this paper cites.
Y. LeCun, J. Denker, and S. Solla, “Optimal brain damage,” Advances in neural information processing systems , vol. 2, 1989
1989
Earlier work this paper cites.
D. P. Bertsekas, “Auction algorithms for network flow problems: A tutorial introduction,” Computational optimization and applications , vol. 1, pp. 7–66, 1992
1992
Earlier work this paper cites.
B. Hassibi, D. G. Stork, and G. J. Wolff, “Optimal brain surgeon and general network pruning,” in IEEE international conference on neural networks . IEEE, 1993, pp. 293–299
1993
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proceedings of the 27th international conference on machine learning (ICML-10) , 2010, pp. 807–814
2010
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
B. Zoph and Q. Le, “Neural architecture search with reinforcement learning,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
NVIDIA, “cublas: Basic linear algebra on nvidia gpus,” [Online], 2017, https://developer.nvidia.com/cublas
2017
Earlier work this paper cites.
——, “Cutlass: Cuda templates for linear algebra subroutines,” [Online], 2017, https://github.com/NVIDIA/cutlass
2017
Earlier work this paper cites.
NVIDIA, “Fastertransformer: About transformer related optimization, including bert, gpt,” [Online], 2017, https://github.com/NVIDIA/FasterTransformer
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
D. C. Mocanu, E. Mocanu, P. Stone, P. H. Nguyen, M. Gibescu, and A. Liotta, “Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science,” Nature communications , vol. 9, no. 1, p. 2383, 2018
2018
Earlier work this paper cites.
M. Stern, N. Shazeer, and J. Uszkoreit, “Blockwise parallel decoding for deep autoregressive models,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Lee, Y. Lee, J. Kim, A. Kosiorek, S. Choi, and Y. W. Teh, “Set transformer: A framework for attention-based permutation-invariant neural networks,” in International conference on machine learning . PMLR, 2019, pp. 3744–3753
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Tillet, H. T. Kung, and D. Cox, “Triton: an intermediate language and compiler for tiled neural network computations,” in Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages , 2019, pp. 10–19
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 9459–9474, 2020
2020
Earlier work this paper cites.
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré, “Hippo: Recurrent memory with optimal polynomial projections,” Advances in neural information processing systems , vol. 33, pp. 1474–1487, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
G. I. Winata, S. Cahyawijaya, Z. Lin, Z. Liu, and P. Fung, “Lightweight and efficient end-to-end speech recognition using low-rank transformer,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6144–6148
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International conference on machine learning . PMLR, 2020, pp. 5156–5165
2020
Earlier work this paper cites.
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser et al. , “Rethinking attention with performers,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Z. Dai, G. Lai, Y. Yang, and Q. Le, “Funnel-transformer: Filtering out sequential redundancy for efficient language processing,” Advances in neural information processing systems , vol. 33, pp. 4271–4282, 2020
2020
Earlier work this paper cites.
W. Liu, P. Zhou, Z. Wang, Z. Zhao, H. Deng, and Q. Ju, “Fastbert: a self-distilling bert with adaptive inference time,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 6035–6044
2020
Earlier work this paper cites.
J. Xin, R. Tang, J. Lee, Y. Yu, and J. Lin, “Deebert: Dynamic early exiting for accelerating bert inference,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 2246–2251
2020
Earlier work this paper cites.
W. Zhou, C. Xu, T. Ge, J. McAuley, K. Xu, and F. Wei, “Bert loses patience: Fast and robust inference with early exit,” Advances in Neural Information Processing Systems , vol. 33, pp. 18 330–18 341, 2020
2020
Earlier work this paper cites.
L. Hou, Z. Huang, L. Shang, X. Jiang, X. Chen, and Q. Liu, “Dynabert: Dynamic bert with adaptive width and depth,” Advances in Neural Information Processing Systems , vol. 33, pp. 9782–9793, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang et al. , “Big bird: Transformers for longer sequences,” Advances in neural information processing systems , vol. 33, pp. 17 283–17 297, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Tay, D. Bahri, L. Yang, D. Metzler, and D.-C. Juan, “Sparse sinkhorn attention,” in International Conference on Machine Learning . PMLR, 2020, pp. 9438–9447
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 4582–4597
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré, “Combining recurrent, convolutional, and continuous-time models with linear state space layers,” Advances in neural information processing systems , vol. 34, pp. 572–585, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
X. Ma, X. Kong, S. Wang, C. Zhou, J. May, H. Ma, and L. Zettlemoyer, “Luna: Linear unified nested attention,” Advances in Neural Information Processing Systems , vol. 34, pp. 2441–2453, 2021
2021
Earlier work this paper cites.
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer, “Base layers: Simplifying training of large, sparse models,” in International Conference on Machine Learning . PMLR, 2021, pp. 6265–6274
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, pp. 97–110
2021
Earlier work this paper cites.
A. Roy, M. Saffar, A. Vaswani, and D. Grangier, “Efficient content-based sparse attention with routing transformers,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 53–68, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y. Han, G. Huang, S. Song, L. Yang, H. Wang, and Y. Wang, “Dynamic neural networks: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 11, pp. 7436–7456, 2021
2021
Earlier work this paper cites.
T. J. Ham, Y. Lee, S. H. Seo, S. Kim, H. Choi, S. J. Jun, and J. W. Lee, “Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks,” in ACM/IEEE 48th Annual International Symposium on Computer Architecture , 2021, pp. 692–705
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Gupta, A. Gu, and J. Berant, “Diagonal state spaces are as effective as structured state spaces,” Advances in Neural Information Processing Systems , vol. 35, pp. 22 982–22 994, 2022
2022
Earlier work this paper cites.
A. Gu, K. Goel, A. Gupta, and C. Ré, “On the parameterization and initialization of diagonal state space models,” Advances in Neural Information Processing Systems , vol. 35, pp. 35 971–35 983, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. Smith, and L. Kong, “Random feature attention,” in International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” The Journal of Machine Learning Research , vol. 23, no. 1, pp. 5232–5270, 2022
2022
Earlier work this paper cites.
Z. Zhang, Y. Lin, Z. Liu, P. Li, M. Sun, and J. Zhou, “Moefication: Transformer feed-forward layers are mixtures of experts,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 877–890
2022
Earlier work this paper cites.
Z.-F. Gao, P. Liu, W. X. Zhao, Z.-Y. Lu, and J.-R. Wen, “Parameter-efficient mixture-of-experts architecture for pre-trained language models,” in Proceedings of the 29th International Conference on Computational Linguistics , 2022, pp. 3263–3273
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Q. V. Le, J. Laudon et al. , “Mixture-of-experts with expert choice routing,” Advances in Neural Information Processing Systems , vol. 35, pp. 7103–7114, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
D. Dai, L. Dong, S. Ma, B. Zheng, Z. Sui, B. Chang, and F. Wei, “Stablemoe: Stable routing strategy for mixture of experts,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 7085–7095
2022
Earlier work this paper cites.
T. Chen, Z. Zhang, A. K. JAISWAL, S. Liu, and Z. Wang, “Sparse moe as the new dropout: Scaling dense and self-slimmable transformers,” in The Eleventh International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat et al. , “Glam: Efficient scaling of language models with mixture-of-experts,” in International Conference on Machine Learning . PMLR, 2022, pp. 5547–5569
2022
Earlier work this paper cites.
W. Hua, Z. Dai, H. Liu, and Q. Le, “Transformer quality in linear time,” in International Conference on Machine Learning . PMLR, 2022, pp. 9099–9117
2022
Earlier work this paper cites.
T. Schuster, A. Fisch, J. Gupta, M. Dehghani, D. Bahri, V. Tran, Y. Tay, and D. Metzler, “Confident adaptive language modeling,” Advances in Neural Information Processing Systems , vol. 35, pp. 17 456–17 472, 2022
2022
Earlier work this paper cites.
J. Kong, J. Wang, L.-C. Yu, and X. Zhang, “Accelerating inference for pretrained language models by unified multi-perspective early exiting,” in Proceedings of the 29th International Conference on Computational Linguistics , 2022, pp. 4677–4686
2022
Earlier work this paper cites.
T. Sun, X. Liu, W. Zhu, Z. Geng, L. Wu, Y. He, Y. Ni, G. Xie, X.-J. Huang, and X. Qiu, “A simple hash-based early exiting approach for language understanding and generation,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 2409–2421
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Javaheripi, G. de Rosa, S. Mukherjee, S. Shah, T. Religa, C. C. Teodoro Mendes, S. Bubeck, F. Koushanfar, and D. Dey, “Litetransformersearch: Training-free neural architecture search for efficient language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 254–24 267, 2022
2022
Earlier work this paper cites.
D. D. Xu, S. Mukherjee, X. Liu, D. Dey, W. Wang, X. Zhang, A. Awadallah, and J. Gao, “Few-shot task-agnostic neural architecture search for distilling large language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 28 644–28 656, 2022
2022
Earlier work this paper cites.
E. Kurtic, D. Campos, T. Nguyen, E. Frantar, M. Kurtz, B. Fineran, M. Goin, and D. Alistarh, “The optimal bert surgeon: Scalable and accurate second-order pruning for large language models,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 4163–4181
2022
Earlier work this paper cites.
W. Kwon, S. Kim, M. W. Mahoney, J. Hassoun, K. Keutzer, and A. Gholami, “A fast post-training pruning framework for transformers,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 101–24 116, 2022
2022
Earlier work this paper cites.
Q. Zhang, S. Zuo, C. Liang, A. Bukharin, P. He, W. Chen, and T. Zhao, “Platon: Pruning large transformer models with upper confidence bound of weight importance,” in International Conference on Machine Learning . PMLR, 2022, pp. 26 809–26 823
2022
Earlier work this paper cites.
M. Xia, Z. Zhong, and D. Chen, “Structured pruning learns compact and accurate models,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 1513–1528
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Yao, R. Y. Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” in Advances in Neural Information Processing Systems , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. Frantar and D. Alistarh, “Optimal brain compression: A framework for accurate post-training quantization and pruning,” in Advances in Neural Information Processing Systems , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022
2022
Earlier work this paper cites.
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al. , “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2022, pp. 1–15
2022
Earlier work this paper cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “Glm: General language model pretraining with autoregressive blank infilling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 320–335
2022
Earlier work this paper cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for transformer-based generative models,” in Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation , 2022, pp. 521–538
2022
Earlier work this paper cites.
H. Fan, T. Chau, S. I. Venieris, R. Lee, A. Kouris, W. Luk, N. D. Lane, and M. S. Abdelfattah, “Adaptable butterfly accelerator for attention-based nns via hardware and algorithm co-design,” in IEEE/ACM International Symposium on Microarchitecture , 2022, pp. 599–615
2022
Earlier work this paper cites.
S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y. Kim, “Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation,” in IEEE Hot Chips 34 Symposium , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez et al. , “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” See https://vicuna. lmsys. org (accessed 14 April 2023) , 2023
2023
Earlier work this paper cites.
D. Li, R. Shao, A. Xie, Y. Sheng, L. Zheng, J. Gonzalez, I. Stoica, X. Ma, and H. Zhang, “How long can context length of open-source llms truly promise?” in NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Yang, H. Jin, R. Tang, X. Han, Q. Feng, H. Jiang, S. Zhong, B. Yin, and X. Hu, “Harnessing the power of llms in practice: A survey on chatgpt and beyond,” ACM Transactions on Knowledge Discovery from Data , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 274–19 286
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Shi, S. Min, M. Yasunaga, M. Seo, R. James, M. Lewis, L. Zettlemoyer, and W. tau Yih, “Replug: Retrieval-augmented black-box language models,” 2023
2023
Cited alongside, same era.
A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-rag: Learning to retrieve, generate, and critique through self-reflection,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Fu, P. Bailis, I. Stoica, and H. Zhang, “Breaking the sequential dependency of llm inference using lookahead decoding,” November 2023. [Online]. Available: https://lmsys.org/blog/2023-11-21-lookahead-decoding/
2023
Later among the works it cites.
Y. Li, C. Zhang, and H. Zhang, “Eagle: Lossless acceleration of llm decoding by feature extrapolation,” December 2023. [Online]. Available: https://sites.google.com/view/eagle-llm
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
T. Gale, D. Narayanan, C. Young, and M. Zaharia, “Megablocks: Efficient sparse training with mixture-of-experts,” in Proceedings of Machine Learning and Systems (MLSys) , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Zhai, C. Jiang, L. Wang, X. Jia, S. Zhang, Z. Chen, X. Liu, and Y. Zhu, “Bytetransformer: A high-performance transformer boosted for variable-length inputs,” in 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2023, pp. 344–355
2023
Later among the works it cites.
T. Dao, D. Haziza, F. Massa, and G. Sizov, “Flash-decoding for long-context inference,” [Online], 2023, https://crfm.stanford.edu/2023/10/12/flashdecoding.html
2023
Later among the works it cites.
2023
Later among the works it cites.
Sensetime, “Openppl: A high-performance deep learning inference platform,” [Online], 2023, https://openppl.ai/home
2023
Later among the works it cites.
S. Wang, “Fastgemv: High-speed gemv kernels,” [Online], 2023, https://github.com/wangsiping97/FastGEMV
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Jin, C.-F. Wu, D. Brooks, and G.-Y. Wei, “S 3
2023
Later among the works it cites.
Y. Qin, Y. Wang, D. Deng, Z. Zhao, X. Yang, L. Liu, S. Wei, Y. Hu, and S. Yin, “Fact: Ffn-attention co-optimized transformer architecture with eager correlation prediction,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023, pp. 1–14
2023
Later among the works it cites.
2023
Later among the works it cites.
S. teams, “Sharegpt,” 2023. [Online]. Available: https://sharegpt.com/
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff et al. , “Pythia: A suite for analyzing large language models across training and scaling,” in International Conference on Machine Learning . PMLR, 2023, pp. 2397–2430
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Geng and H. Liu, “Openllama: An open reproduction of llama,” May 2023. [Online]. Available: https://github.com/openlm-research/open_llama
2023
Later among the works it cites.
M. team, “MLC-LLM,” 2023. [Online]. Available: https://github.com/mlc-ai/mlc-llm
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao, “Medusa: Simple llm inference acceleration framework with multiple decoding heads,” 2024
2024
Closest in time.
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk et al. , “Graph of thoughts: Solving elaborate problems with large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 16, 2024, pp. 17 682–17 690
2024
Closest in time.
2024
Closest in time.
J. Pilault, M. Fathi, O. Firat, C. Pal, P.-L. Bacon, and R. Goroshin, “Block-state transformers,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
AI21, “Jamba: Ai21’s groundbreaking ssm-transformer model,” March 2024. [Online]. Available: https://www.ai21.com/blog/announcing-jamba
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett et al. , “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
X. Ma, G. Fang, and X. Wang, “Llm-pruner: On the structural pruning of large language models,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
E. Kurtić, E. Frantar, and D. Alistarh, “Ziplm: Inference-aware structured pruning of language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
P. Dong, L. Li, Z. Tang, X. Liu, X. Pan, Q. Wang, and X. Chu, “Pruner-zero: Evolving symbolic pruning metric from scratch for large language models,” in International Conference on Machine Learning (ICML) , 2024
2024
Closest in time.
Y. Zhang, L. Zhao, M. Lin, S. Yunyun, Y. Yao, X. Han, J. Tanner, S. Liu, and R. Ji, “Dynamic sparse no training: Training-free fine-tuning for sparse llms,” in International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
W. Huang, Y. Liu, H. Qin, Y. Li, S. Zhang, X. Liu, M. Magno, and X. Qi, “Billm: Pushing the limit of post-training quantization for llms,” 2024
2024
Closest in time.
2024
Closest in time.
Y. Ma, H. Li, X. Zheng, F. Ling, X. Xiao, R. Wang, S. Wen, F. Chao, and R. Ji, “Affinequant: Affine transformation quantization for large language models,” in International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
InternLM, “Lmdeploy,” 2024. [Online]. Available: https://github.com/InternLM/lmdeploy
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
ggerganov, “Inference of meta’s llama model (and others) in pure c/c++,” 2024. [Online]. Available: https://github.com/ggerganov/llama.cpp
2024
Closest in time.
2024
Closest in time.
K. Hong, G. Dai, J. Xu, Q. Mao, X. Li, J. Liu, K. Chen, Y. Dong, and Y. Wang, “Flashdecoding++: Faster large language model inference on gpus,” 2024
2024
Closest in time.
HuggingFace, “Transformers: State-of-the-art machine learning for pytorch, tensorflow, and jax.” [Online], 2024, https://github.com/huggingface/transformers
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
ModelTC, “Lightllm,” February 2024. [Online]. Available: https://github.com/ModelTC/lightllm/
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Ye, “flashinfer,” March 2024. [Online]. Available: https://github.com/flashinfer-ai/flashinfer
2024
Closest in time.
H. Oh, K. Kim, J. Kim, S. Kim, J. Lee, D.-s. Chang, and J. Seo, “Exegpt: Constraint-aware resource scheduling for llm inference,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2024, pp. 369–384
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
“Minicpm: Unveiling the potential of end-side large language models,” 2024
2024
Closest in time.
2024
Closest in time.
Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang, “A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,” High-Confidence Computing , p. 100211, 2024
2024
Closest in time.