Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have revolutionized a wide range of domains such as natural language processing, computer vision, and multi-modal tasks due to their ability to comprehend context and perform logical reasoning.
1904
Earlier work this paper cites.
1906
Earlier work this paper cites.
1911
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , P. Isabelle, E. Charniak, and D. Lin, Eds. Philadelphia, Pennsylvania, USA: Association for Computational Linguistics, Jul. 2002, pp. 311–318. [Online]. Available: https://aclanthology.org/P02-1040
2002
Earlier work this paper cites.
S. Podlipnig and L. Böszörmenyi, “A survey of web cache replacement strategies,” ACM Computing Surveys (CSUR) , vol. 35, no. 4, pp. 374–398, 2003
2003
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summarization Branches Out . Barcelona, Spain: Association for Computational Linguistics, Jul. 2004, pp. 74–81. [Online]. Available: https://aclanthology.org/W04-1013
2004
Earlier work this paper cites.
H. Jegou, M. Douze, and C. Schmid, “Product quantization for nearest neighbor search,” IEEE transactions on pattern analysis and machine intelligence , vol. 33, no. 1, pp. 117–128, 2010
2010
Earlier work this paper cites.
M. Denkowski and A. Lavie, “Meteor 1.3: Automatic metric for reliable optimization and evaluation of machine translation systems,” in Proceedings of the Sixth Workshop on Statistical Machine Translation , C. Callison-Burch, P. Koehn, C. Monz, and O. F. Zaidan, Eds. Edinburgh, Scotland: Association for Computational Linguistics, Jul. 2011, pp. 85–91. [Online]. Available: https://aclanthology.org/W11-2107
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
G. Gracioli, A. Alhammad, R. Mancuso, A. A. Fröhlich, and R. Pellizzoni, “A survey on cache management mechanisms for real-time embedded systems,” ACM Computing Surveys (CSUR) , vol. 48, no. 2, pp. 1–36, 2015
2015
Earlier work this paper cites.
V. Kuleshov, A. Chaganty, and P. Liang, “Tensor factorization via matrix factorization,” in Artificial Intelligence and Statistics . PMLR, 2015, pp. 507–516
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Lin, S. Talathi, and S. Annapureddy, “Fixed point quantization of deep convolutional networks,” in International conference on machine learning . PMLR, 2016, pp. 2849–2858
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
J. Cheng, J. Wu, C. Leng, Y. Wang, and Q. Hu, “Quantized cnn: A unified approach to accelerate and compress convolutional networks,” IEEE transactions on neural networks and learning systems , vol. 29, no. 10, pp. 4730–4743, 2017
2017
Earlier work this paper cites.
P. Zhou, C. Lu, Z. Lin, and C. Zhang, “Tensor factorization for low-rank tensor completion,” IEEE Transactions on Image Processing , vol. 27, no. 3, pp. 1152–1163, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video question answering via gradually refined attention over appearance and motion,” in ACM Multimedia , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
A. F. Agarap, “Deep learning using rectified linear units,” arXiv preprint arXiv:1803.08375 , 2018
2018
Earlier work this paper cites.
Y. Zhou, S.-M. Moosavi-Dezfooli, N.-M. Cheung, and P. Frossard, “Adaptive quantization for deep neural network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
P. Jiang and G. Agrawal, “A linear speedup analysis of distributed deep learning with sparse and quantized communication,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
Y. Matsui, Y. Uchida, H. Jégou, and S. Satoh, “A survey of product quantization,” ITE Transactions on Media Technology and Applications , vol. 6, no. 1, pp. 2–10, 2018
2018
Earlier work this paper cites.
O. A. Malik and S. Becker, “Low-rank tucker decomposition of large tensors using tensorsketch,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Post, “A call for clarity in reporting BLEU scores,” in Proceedings of the Third Conference on Machine Translation: Research Papers , O. Bojar, R. Chatterjee, C. Federmann, M. Fishel, Y. Graham, B. Haddow, M. Huck, A. J. Yepes, P. Koehn, C. Monz, M. Negri, A. Névéol, M. Neves, M. Post, L. Specia, M. Turchi, and K. Verspoor, Eds. Brussels, Belgium: Association for Computational Linguistics, Oct. 2018, pp. 186–191. [Online]. Available: https://aclanthology.org/W18-6319
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
A. Kwasniewska, M. Szankin, M. Ozga, J. Wolfe, A. Das, A. Zajac, J. Ruminski, and P. Rad, “Deep learning optimization for edge devices: Analysis of training quantization parameters,” in IECON 2019-45th Annual Conference of the IEEE Industrial Electronics Society , vol. 1. IEEE, 2019, pp. 96–101
2019
Earlier work this paper cites.
E. Sharma, C. Li, and L. Wang, “BIGPATENT: A large-scale dataset for abstractive and coherent summarization,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , A. Korhonen, D. Traum, and L. Màrquez, Eds. Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 2204–2213. [Online]. Available: https://aclanthology.org/P19-1212/
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
Z. Lin, C. Li, Y. Miao, Y. Liu, and Y. Xu, “Pagraph: Scaling gnn training on large graphs via computation-aware caching,” in Proceedings of the 11th ACM Symposium on Cloud Computing , 2020, pp. 401–415
2020
Earlier work this paper cites.
J. W. Rae, A. Potapenko, S. M. Jayakumar, C. Hillier, and T. P. Lillicrap, “Compressive transformers for long-range sequence modelling,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net, 2020. [Online]. Available: https://openreview.net/forum?id=SylKikSYDH
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International conference on machine learning . PMLR, 2020, pp. 5156–5165
2020
Earlier work this paper cites.
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser et al. , “Rethinking attention with performers,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
X. Ho, A.-K. Duong Nguyen, S. Sugawara, and A. Aizawa, “Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps,” in Proceedings of the 28th International Conference on Computational Linguistics , D. Scott, N. Bel, and C. Zong, Eds. Barcelona, Spain (Online): International Committee on Computational Linguistics, Dec. 2020, pp. 6609–6625. [Online]. Available: https://aclanthology.org/2020.coling-main.580/
2020
Earlier work this paper cites.
H. Li and L. Chen, “Cache-based GNN system for dynamic graphs,” in CIKM ’21: The 30th ACM International Conference on Information and Knowledge Management, Virtual Event, Queensland, Australia, November 1 - 5, 2021 . ACM, 2021, pp. 937–946. [Online]. Available: https://doi.org/10.1145/3459637.3482237
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Gu, I. Johnson, K. Goel, K. K. Saab, T. Dao, A. Rudra, and C. Re, “Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space Layers,” in Advances in Neural Information Processing Systems , Nov. 2021. [Online]. Available: https://openreview.net/forum?id=yWd42CWN3c
2021
Earlier work this paper cites.
W. Dong and K. Yi, “Residual sensitivity for differentially private multi-way joins,” in Proceedings of the 2021 International Conference on Management of Data , 2021, pp. 432–444
2021
Earlier work this paper cites.
J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2021, pp. 9777–9786
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Angelidis, R. K. Amplayo, Y. Suhara, X. Wang, and M. Lapata, “Extractive opinion summarization in quantized transformer spaces,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 277–293, 03 2021. [Online]. Available: https://doi.org/10.1162/tacl_a_00366
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “GLM: general language model pretraining with autoregressive blank infilling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Association for Computational Linguistics, 2022, pp. 320–335. [Online]. Available: https://doi.org/10.18653/v1/2022.acl-long.26
2022
Earlier work this paper cites.
Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 168–27 183, 2022
2022
Earlier work this paper cites.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and memory-efficient exact attention with IO-awareness,” in Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,” Advances in Neural Information Processing Systems , vol. 35, pp. 30 318–30 332, 2022
2022
Earlier work this paper cites.
X. Wei, Y. Zhang, X. Zhang, R. Gong, S. Zhang, Q. Zhang, F. Yu, and X. Liu, “Outlier suppression: Pushing the limit of low-bit transformer language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 17 402–17 414, 2022
2022
Earlier work this paper cites.
W. Hua, Z. Dai, H. Liu, and Q. V. Le, “Transformer Quality in Linear Time,” in International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA , ser. Proceedings of Machine Learning Research, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvári, G. Niu, and S. Sabato, Eds., vol. 162. PMLR, 2022, pp. 9099–9117. [Online]. Available: https://proceedings.mlr.press/v162/hua22a.html
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Gu, K. Goel, A. Gupta, and C. Ré, “On the Parameterization and Initialization of Diagonal State Space Models,” in Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022 , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds. arXiv, 2022
2022
Earlier work this paper cites.
G. Yu, J. S. Jeong, G. Kim, S. Kim, and B. Chun, “Orca: A distributed serving system for transformer-based generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2022, Carlsbad, CA, USA, July 11-13, 2022 , M. K. Aguilera and H. Weatherspoon, Eds. USENIX Association, 2022, pp. 521–538. [Online]. Available: https://www.usenix.org/conference/osdi22/presentation/yu
2022
Earlier work this paper cites.
Y. Zhao and J. Chen, “A survey on differential privacy for unstructured data content,” ACM Computing Surveys (CSUR) , vol. 54, no. 10s, pp. 1–28, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Y. Pang, A. Parrish, N. Joshi, N. Nangia, J. Phang, A. Chen, V. Padmakumar, J. Ma, J. Thompson, H. He, and S. Bowman, “QuALITY: Question answering with long input texts, yes!” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Seattle, United States: Association for Computational Linguistics, Jul. 2022, pp. 5336–5358. [Online]. Available: https://aclanthology.org/2022.naacl-main.391
2022
Earlier work this paper cites.
M. U. Hadi, R. Qureshi, A. Shah, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili et al. , “A survey on large language models: Applications, challenges, limitations, and practical usage,” Authorea Preprints , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023. [Online]. Available: http://papers.nips.cc/paper_files/paper/2023/hash/6dcf277ea32ce3288914faf369fe6de0-Abstract-Conference.html
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Wu, W. Gan, Z. Chen, S. Wan, and P. S. Yu, “Multimodal large language models: A survey,” in 2023 IEEE International Conference on Big Data (BigData) . IEEE, 2023, pp. 2247–2256
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Yang, X. Jia, H. Li, and J. Yan, “Llm4drive: A survey of large language models for autonomous driving,” in NeurIPS 2024 Workshop on Open-World Agents , 2023
2023
Earlier work this paper cites.
J. Qiu, L. Li, J. Sun, J. Peng, P. Shi, R. Zhang, Y. Dong, K. Lam, F. P.-W. Lo, B. Xiao et al. , “Large ai models in health informatics: Applications, challenges, and the future,” IEEE Journal of Biomedical and Health Informatics , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Zeng, X. Liu et al. , “GLM-130B: an open bilingual pre-trained model,” in The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023. [Online]. Available: https://openreview.net/forum?id=-Aw0rrrPUF
2023
Earlier work this paper cites.
Y. Li, Y. Shen, L. Chen, and M. Yuan, “Orca: Scalable temporal graph neural network training with theoretical guarantees,” Proc. ACM Manag. Data , vol. 1, no. 1, pp. 52:1–52:27, 2023. [Online]. Available: https://doi.org/10.1145/3588737
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 274–19 286
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Xi, C. Li, J. Chen, and J. Zhu, “Training transformers with 4-bit integers,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023. [Online]. Available: http://papers.nips.cc/paper_files/paper/2023/hash/99fc8bc48b917c301a80cb74d91c0c06-Abstract-Conference.html
2023
Earlier work this paper cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, 2023, pp. 38 087–38 099. [Online]. Available: https://proceedings.mlr.press/v202/xiao23c.html
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single GPU,” in International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, 2023, pp. 31 094–31 116. [Online]. Available: https://proceedings.mlr.press/v202/sheng23a.html
2023
Earlier work this paper cites.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. W. Barrett, Z. Wang, and B. Chen, “H2O: heavy-hitter oracle for efficient generative inference of large language models,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023. [Online]. Available: http://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b15572587b78ecfcebb2827f8-Abstract-Conference.html
2023
Earlier work this paper cites.
Z. Liu, A. Desai, F. Liao, W. Wang, V. Xie, Z. Xu, A. Kyrillidis, and A. Shrivastava, “Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , 2023
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023 . ACM, 2023, pp. 611–626
2023
Earlier work this paper cites.
J. Mu, X. Li, and N. D. Goodman, “Learning to compress prompts with gist tokens,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023. [Online]. Available: http://papers.nips.cc/paper_files/paper/2023/hash/3d77c6dcc7f143aa2154e7f4d5e22d68-Abstract-Conference.html
2023
Earlier work this paper cites.
Y. Bondarenko, M. Nagel, and T. Blankevoort, “Quantizable transformers: Removing outliers by helping attention heads do nothing,” Advances in Neural Information Processing Systems , vol. 36, pp. 75 067–75 096, 2023
2023
Earlier work this paper cites.
D. Zhou, K. Wang, J. Gu, X. Peng, D. Lian, Y. Zhang, Y. You, and J. Feng, “Dataset quantization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 17 205–17 216
2023
Cited alongside, same era.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett et al. , “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, pp. 34 661–34 710, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Dong, Q. Luo, and K. Yi, “Continual observation under user-level differential privacy,” in 2023 IEEE Symposium on Security and Privacy (SP) . IEEE, 2023, pp. 2190–2207
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
M. Adnan, A. Arunkumar, G. Jain, P. J. Nair, I. Soloveychik, and P. Kamath, “Keyformer: KV cache reduction through key tokens selection for efficient generative inference,” in Proceedings of the Seventh Annual Conference on Machine Learning and Systems, MLSys 2024, Santa Clara, CA, USA, May 13-16, 2024 . mlsys.org, 2024
2024
Closest in time.
2024
Closest in time.
S. Ge, Y. Zhang, L. Liu, M. Zhang, J. Han, and J. Gao, “Model tells you what to discard: Adaptive KV cache compression for llms,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. [Online]. Available: https://openreview.net/forum?id=NG7sS51zVF
2024
Closest in time.
C. Han, Q. Wang, H. Peng, W. Xiong, Y. Chen, H. Ji, and S. Wang, “Lm-infinite: Zero-shot extreme length generalization for large language models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024 . Association for Computational Linguistics, 2024, pp. 3991–4008
2024
Closest in time.
L. Ribar, I. Chelombiev, L. Hudlass-Galley, C. Blake, C. Luschi, and D. Orr, “Sparq attention: Bandwidth-efficient LLM inference,” in Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Anonymous, “CAKE: Cascading and adaptive KV cache eviction with layer preferences,” in Submitted to The Thirteenth International Conference on Learning Representations , 2024, under review. [Online]. Available: https://openreview.net/forum?id=EQgEMAD4kv
2024
Closest in time.
2024
Closest in time.
T. Dao, “FlashAttention-2: Faster attention with better parallelism and work partitioning,” in International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Liu, H. Li, Y. Cheng, S. Ray, Y. Huang, Q. Zhang, K. Du, J. Yao, S. Lu, G. Ananthanarayanan et al. , “Cachegen: Kv cache compression and streaming for fast large language model serving,” in Proceedings of the ACM SIGCOMM 2024 Conference , 2024, pp. 38–56
2024
Closest in time.
Y. Zhao, C.-Y. Lin, K. Zhu, Z. Ye, L. Chen, S. Zheng, L. Ceze, A. Krishnamurthy, T. Chen, and B. Kasikci, “Atom: Low-bit quantization for efficient and accurate llm serving,” Proceedings of Machine Learning and Systems , vol. 6, pp. 196–209, 2024
2024
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Xu and Z. Lin, “MixCon: A Hybrid Architecture for Efficient and Adaptive Sequence Modeling,” in Frontiers in Artificial Intelligence and Applications , U. Endriss, F. S. Melo, K. Bach, A. Bugarín-Diz, J. M. Alonso-Moral, S. Barro, and F. Heintz, Eds. IOS Press, Oct. 2024. [Online]. Available: https://ebooks.iospress.nl/doi/10.3233/FAIA240593
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Yen, T. Gao, and D. Chen, “Long-Context Language Modeling with Parallel Context Encoding,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024 , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Association for Computational Linguistics, 2024, pp. 2588–2610. [Online]. Available: https://doi.org/10.18653/v1/2024.acl-long.142
2024
Closest in time.
J. Monteiro, É. Marcotte, P.-A. Noël, V. Zantedeschi, D. Vázquez, N. Chapados, C. Pal, and P. Taslakian, “XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference,” in Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024 , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Association for Computational Linguistics, 2024, pp. 15 284–15 302. [Online]. Available: https://aclanthology.org/2024.findings-emnlp.896
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Wu and K. Tu, “Layer-Condensed KV Cache for Efficient Inference of Large Language Models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024 , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Association for Computational Linguistics, 2024, pp. 11 175–11 188. [Online]. Available: https://doi.org/10.18653/v1/2024.acl-long.602
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Chen, C. Zhang, X. Gao, R. D. Mullins, G. A. Constantinides, and Y. Zhao, “Optimised Grouped-Query Attention Mechanism for Transformers,” in Workshop on Efficient Systems for Foundation Models II @ ICML2024 , Jul. 2024. [Online]. Available: https://openreview.net/forum?id=13MMghY6Kh
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024 , A. Gavrilovska and D. B. Terry, Eds. USENIX Association, 2024, pp. 193–210. [Online]. Available: https://www.usenix.org/conference/osdi24/presentation/zhong-yinmin
2024
Closest in time.
O. Guldogan, J. Kunde, K. Lee, and R. Pedarsani, “Multi-bin batching for increasing LLM inference throughput,” 2024. [Online]. Available: https://openreview.net/forum?id=WVmarX0RNd
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Gao, X. Zhang, Y. Shen, and L. Chen, “Apt-serve: Adaptive request scheduling on hybrid cache for scalable llm inference serving,” Proceedings of the ACM on Management of Data , vol. 3, no. 3, p. 1–28, Jun. 2025. [Online]. Available: http://dx.doi.org/10.1145/3725394
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.