Fetching the paper…
Reading the bibliography…
The wide deployment of Large Language Models (LLMs) has given rise to strong demands for optimizing their inference performance.
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
H. Naghibijouybari, A. Neupane, Z. Qian, and N. Abu-Ghazaleh, “Rendered insecure: Gpu side channel attacks are practical,” in Proceedings of the 2018 ACM SIGSAC conference on computer and communications security , 2018, pp. 2139–2153
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Schwarz, M. Schwarzl, M. Lipp, J. Masters, and D. Gruss, “Netspectre: Read arbitrary memory over network,” in Computer Security–ESORICS 2019: 24th European Symposium on Research in Computer Security, Luxembourg, September 23–27, 2019, Proceedings, Part I 24 . Springer, 2019, pp. 279–299
2019
Earlier work this paper cites.
“Medquad,” https://huggingface.co/datasets/lavita/MedQuAD , 2019
2019
Earlier work this paper cites.
M. Kurth, B. Gras, D. Andriesse, C. Giuffrida, H. Bos, and K. Razavi, “Netcat: Practical cache attacks from the network,” in 2020 IEEE Symposium on Security and Privacy (SP) . IEEE, 2020, pp. 20–38
2020
Earlier work this paper cites.
M. Yan, C. W. Fletcher, and J. Torrellas, “Cache telepathy: Leveraging shared resource attacks to learn { \{ DNN } \} architectures,” in 29th USENIX Security Symposium (USENIX Security 20) , 2020, pp. 2003–2020
2020
Earlier work this paper cites.
C. Gongye, Y. Fei, and T. Wahl, “Reverse-engineering deep neural networks using floating-point timing side-channels,” in 2020 57th ACM/IEEE Design Automation Conference (DAC) . IEEE, 2020, pp. 1–6
2020
Earlier work this paper cites.
J. Wei, Y. Zhang, Z. Zhou, Z. Li, and M. A. Al Faruque, “Leaky dnn: Stealing deep-learning model secret with gpu context-switching side-channel,” in 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) . IEEE, 2020, pp. 125–137
2020
Earlier work this paper cites.
Y. Zhang, R. Yasaei, H. Chen, Z. Li, and M. A. Al Faruque, “Stealing neural network structure through remote fpga side-channel analysis,” IEEE Transactions on Information Forensics and Security , vol. 16, pp. 4377–4388, 2021
2021
Earlier work this paper cites.
Y. Zhu, Y. Cheng, H. Zhou, and Y. Lu, “Hermes attack: Steal DNN models with lossless inference accuracy,” in 30th USENIX Security Symposium (USENIX Security 21) , 2021
2021
Earlier work this paper cites.
Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 168–27 183, 2022
2022
Earlier work this paper cites.
X. Wei, Y. Zhang, X. Zhang, R. Gong, S. Zhang, Q. Zhang, F. Yu, and X. Liu, “Outlier suppression: Pushing the limit of low-bit transformer language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 17 402–17 414, 2022
2022
Earlier work this paper cites.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022
2022
Earlier work this paper cites.
“Presidio,” https://github.com/microsoft/presidio , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. S. Rakin, M. H. I. Chowdhuryy, F. Yao, and D. Fan, “Deepsteal: Advanced model extractions leveraging efficient weight stealing in memories,” in 2022 IEEE symposium on security and privacy (SP) . IEEE, 2022, pp. 1157–1174
2022
Earlier work this paper cites.
H. T. Maia, C. Xiao, D. Li, E. Grinspun, and C. Zheng, “Can one hear the shape of a neural network?: Snooping the gpu via magnetic side channel.” in USENIX Security Symposium , 2022, pp. 4383–4400
2022
Earlier work this paper cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 38 087–38 099
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles (SOSP) , 2023, pp. 611–626
2023
Cited alongside, same era.
L. Zheng, L. Yin, Z. Xie, J. Huang, C. Sun, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez et al. , “Efficiently programming large language models using SGLang,” arXiv preprint , 2023
2023
Cited alongside, same era.
F. Bang, “Gptcache: An open-source semantic cache for llm applications enabling faster answers and cost savings,” in Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023) , 2023, pp. 212–218
2023
Cited alongside, same era.
H. Taneja, J. Kim, J. J. Xu, S. Van Schaik, D. Genkin, and Y. Yarom, “Hot pixels: Frequency, power, and temperature attacks on { \{ GPUs } \} and arm { \{ SoCs } \} ,” in 32nd USENIX Security Symposium (USENIX Security 23) , 2023, pp. 6275–6292
“System prompts in large language models,” https://promptengineering.org/system-prompts-in-large-language-models/ , 2024
2024
Closest in time.
“Llamaindex is a data framework for your llm applications,” https://github.com/run-llama/llama_index , 2024
2024
Closest in time.
“Jpmorgan rolls out in-house genai-based chatbot to employees,” https://www.financedirectoreurope.com/news/jpmorgan-rolls-out-ai-based-chatbot/ , 2024
2024
Closest in time.
“Text generation and prompting,” https://platform.openai.com/docs/guides/text?api-mode=responses , 2024
2024
Closest in time.
“System prompt leakage dataset,” https://huggingface.co/datasets/gabrielchua/system-prompt-leakage , 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
“Samsung bans staff’s ai use after spotting chatgpt data leak,” https://www.bloomberg.com/news/articles/2023-05-02/samsung-bans-chatgpt-and-other-generative-ai-use-by-staff-after-leak
2023
Cited alongside, same era.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
“Chatgpt — openai,” https://openai.com/chatgpt/ , 2024
2024
Cited alongside, same era.
“Perplexity ai,” https://www.perplexity.ai/ , 2024
2024
Cited alongside, same era.
S. Group, “flush cache,” https://github.com/sgl-project/sglang/blob/25e5d589e39b3b605296395e4f9c96ec42f09055/python/sglang/srt/server.py#L164 , 2024
2024
Closest in time.
S. Transformers, “all-mpnet-base-v2,” https://huggingface.co/sentence-transformers/all-mpnet-base-v2 , 2024
2024
Closest in time.
“Ai document summarization,” https://www.ibm.com/architectures/hybrid/genai-document-summarization , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
E. Debenedetti, G. Severi, N. Carlini, C. A. Choquette-Choo, M. Jagielski, M. Nasr, E. Wallace, and F. Tramèr, “Privacy side channels in machine learning systems,” in 33rd USENIX Security Symposium , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
G. Wu, Z. Zhang, Y. Zhang, W. Wang, J. Niu, Y. Wu, and Y. Zhang, “I know what you asked: Prompt leakage via kv-cache sharing in multi-tenant llm serving,” in 32nd Annual Network and Distributed System Security Symposium, NDSS 2025, San Diego, California, USA , 2025
2025
Closest in time.
“llm side-channel demo,” https://sites.google.com/view/early-bird-catches-the-leak , 2025
2025
Closest in time.
“Lmcache: Supercharge your llm with the fastest kv cache layer,” https://github.com/LMCache/LMCache , 2025
2025
Closest in time.
2025
Closest in time.
Anonymous, “SemShareKV: Efficient KVCache sharing for semantically similar prompts via token-level LSH matching,” in Submitted to ACL Rolling Review - July 2025 , 2025, under review. [Online]. Available: https://openreview.net/forum?id=B5xRER6OKT
2025
Closest in time.
M. Soleimani, G. Jia, I. Gim, S.-s. Lee, and A. Khandelwal, “Wiretapping llms: Network side-channel attacks on interactive llm services,” Cryptology ePrint Archive , 2025
2025
Closest in time.