Fetching the paper…
Reading the bibliography…
This paper introduces a vision of confidential prompting: securing user prompts from an untrusted, cloud-hosted large language model (LLM) while preserving model confidentiality, output invariance, and compute efficiency.
1909
Earlier work this paper cites.
O. Goldreich, Foundations of cryptography: volume 2, basic applications . Cambridge university press, 2001, vol. 2
2001
Earlier work this paper cites.
F. McKeen, I. Alexandrovich, A. Berenzon, C. V. Rozas, H. Shafi, V. Shanbhogue, and U. R. Savagaonkar, “Innovative instructions and software model for isolated execution.” Hasp@ isca , vol. 10, no. 1, 2013
2013
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems (NeurIPS 2017) , vol. 30, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Demonstrations , W. Ammar, A. Louis, and N. Mostafazadeh, Eds. Association for Computational Linguistics, 2019, pp. 48–53. [Online]. Available: https://doi.org/10.18653/v1/n19-4009
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in Neural Information Processing Systems (NeurIPS 2020) , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Online: Association for Computational Linguistics, Oct. 2020, pp. 38–45. [Online]. Available: https://www.aclweb.org/anthology/2020.emnlp-demos.6
2020
Earlier work this paper cites.
M. Li, Y. Zhang, H. Wang, K. Li, and Y. Cheng, “CIPHERLEAKS: Breaking constant-time cryptography on AMD SEV via the ciphertext side channel,” in 30th USENIX Security Symposium (USENIX Security 21) , 2021, pp. 717–732
2021
Earlier work this paper cites.
Z. Huang, W.-j. Lu, C. Hong, and J. Ding, “Cheetah: Lean and fast secure two-party deep neural network inference,” in 31st USENIX Security Symposium (USENIX Security 22) , 2022, pp. 809–826
2022
Earlier work this paper cites.
M. Hao, H. Li, H. Chen, P. Xing, G. Xu, and T. Zhang, “Iron: Private inference on transformers,” Advances in Neural Information Processing Systems (NeurIPS 2022) , vol. 35, pp. 15 718–15 731, 2022
2022
Earlier work this paper cites.
S. Zhao, M. Li, Y. Zhang, and Z. Lin, “vSGX: Virtualizing SGX enclaves on AMD SEV,” in 2022 IEEE Symposium on Security and Privacy (SP) . IEEE, 2022, pp. 321–336
2022
Earlier work this paper cites.
W. Diffie and M. E. Hellman, “New directions in cryptography,” in Democratizing cryptography: the work of Whitfield Diffie and Martin Hellman , 2022, pp. 365–390
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” Proceedings of Machine Learning and Systems , vol. 5, pp. 606–624, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Qin, Z. Song, W. Zhang, S. Huang, W. Yao, G. Liu, X. Jia, and H. Du, “Protecting encrypted virtual machines from nested page fault controlled channel,” in Proceedings of the Thirteenth ACM Conference on Data and Application Security and Privacy , 2023, pp. 165–175
2023
Earlier work this paper cites.
E. Perot, “Running stable diffusion on gpu with gvisor,” 2025, Accessed: 2025-11-10. [Online]. Available: https://gvisor.dev/blog/2023/06/20/gpu-pytorch-stable-diffusion/
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Y. Akimoto, K. Fukuchi, Y. Akimoto, and J. Sakuma, “Privformer: Privacy-preserving transformer with MPC,” in 2023 IEEE 8th European Symposium on Security and Privacy (EuroS&P) . IEEE, 2023, pp. 392–410
2023
Cited alongside, same era.
2023
Cited alongside, same era.
P. Mai, Y. Yang, R. Yan, R. Ye, and Y. Pang, “ConfusionPrompt: Practical private inference for online large language models,” Available at SSRN 5046754 , 2023
2023
Cited alongside, same era.
——, “Publish, Audit, Attest: How Tinfoil Builds Trust,” 2025, Accessed: 2025-11-10. [Online]. Available: https://tinfoil.sh/blog/2025-01-13-how-tinfoil-builds-trust
2025
Closest in time.
——, “Detailed Attestation Architecture,” 2025, Accessed: 2025-11-10. [Online]. Available: https://docs.tinfoil.sh/verification/attestation-architecture
2025
Closest in time.
Nvidia, “Nvidia confidential computing,” 2023, Accessed: 2025-11-10. [Online]. Available: https://www.nvidia.com/en-us/data-center/solutions/confidential-computing
2025
Closest in time.
——, “NVIDIA Confidential Computing Whitepaper,” 2023, Accessed: 2025-11-10. [Online]. Available: https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/HCC-Whitepaper-v1.0.pdf
2025
Closest in time.
——, “NVIDIA Blackwell Architecture,” 2025, Accessed: 2025-11-10. [Online]. Available: https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “AWQ: Activation-aware weight quantization for LLM compression and acceleration,” in MLSys , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Yang, X. Zhang, Y. Jiang, X. Chen, H. Wang, S. Ji, and Z. Wang, “Prsa: Prompt reverse stealing attacks against large language models,” CoRR , 2024
2024
Cited alongside, same era.
M. Russinovich, “Azure AI Confidential Inferencing: Technical Deep-Dive,” https://techcommunity.microsoft.com/t5/azure-confidential-computing/azure-ai-confidential-inferencing-technical-deep-dive/ba-p/4253150 , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
G. Wu, Z. Zhang, Y. Zhang, W. Wang, J. Niu, Y. Wu, and Y. Zhang, “I know what you asked: Prompt leakage via kv-cache sharing in multi-tenant llm serving,” in Proceedings of the 2025 Network and Distributed System Security (NDSS) Symposium. San Diego, CA, USA , 2025
2025
Closest in time.
Linux developers, “Linux Kernel,” 2025, Accessed: 2025-11-10. [Online]. Available: https://github.com/torvalds/linux
2025
Closest in time.
Nvidia, “NVIDIA Linux open GPU kernel module,” 2025, Accessed: 2025-11-10. [Online]. Available: https://github.com/NVIDIA/open-gpu-kernel-modules
2025
Closest in time.
Nvidia, “CUDA C++ Programming Guide, Interprocess Communication,” 2025, Accessed: 2025-11-10. [Online]. Available: https://docs.nvidia.com/cuda/cuda-c-programming-guide/#interprocess-communication
2025
Closest in time.
Nvidia, “Nvidia MPS,” 2025. [Online]. Available: https://docs.nvidia.com/deploy/mps/index.html
2025
Closest in time.
PyTorch, “PyTorch source code,” 2025, Accessed: 2025-11-10. [Online]. Available: https://github.com/pytorch/pytorch
2025
Closest in time.
VirTEE, “snpguest,” 2025, Accessed: 2025-11-10. [Online]. Available: https://github.com/virtee/snpguest
2025
Closest in time.
Nvidia, “nvTrust,” 2025, Accessed: 2025-11-10. [Online]. Available: https://github.com/NVIDIA/nvtrust
2025
Closest in time.
——, “NVIDIA GPU Attestation Guide,” 2025, Accessed: 2025-11-10. [Online]. Available: https://github.com/NVIDIA/nvtrust/blob/main/guest_tools/README.md
2025
Closest in time.
Github, “Github Actions,” 2025, Accessed: 2025-11-10. [Online]. Available: https://github.com/features/actions
2025
Closest in time.
Nvidia, “Example code for remote attestation of Nvidia GPU,” 2025, Accessed: 2025-11-10. [Online]. Available: https://github.com/NVIDIA/nvtrust/blob/main/guest_tools/attestation_sdk/tests/end_to_end/hardware/test_remote_gpu.py
2025
Closest in time.
PyTorch, “PyTorch CUDA MemPool,” 2025, Accessed: 2025-11-10. [Online]. Available: https://docs.pytorch.org/docs/stable/generated/torch.cuda.memory.MemPool.html
2025
Closest in time.
Nvidia, “CUDA C++ Programming Guide, Shareable Memory Allocations,” 2025, Accessed: 2025-11-10. [Online]. Available: https://docs.nvidia.com/cuda/cuda-c-programming-guide/#shareable-memory-allocations
2025
Closest in time.
Linux developers, “Linux Clone System Call,” 2025, Accessed: 2025-11-10. [Online]. Available: https://man7.org/linux/man-pages/man2/clone.2.html
2025
Closest in time.
Facebook, “Gloo: Collective Communications Library,” 2023, Accessed: 2025-11-10. [Online]. Available: https://github.com/facebookincubator/gloo
2025
Closest in time.
Nvidia, “NVIDIA Collective Communications Library (NCCL),” 2025, Accessed: 2025-11-10. [Online]. Available: https://developer.nvidia.com/nccl
2025
Closest in time.
Y. Yuan, Z. Liu, S. Deng, Y. Chen, S. Wang, Y. Zhang, and Z. Su, “CipherSteal: Stealing input data from TEE-shielded neural networks with ciphertext side channels,” in 2025 IEEE Symposium on Security and Privacy (SP) . IEEE, 2025, pp. 4136–4154
2025
Closest in time.
K. D. Duy, J. Kim, H. Lim, and H. Lee, “INCOGNITOS: A practical unikernel design for full-system obfuscation in confidential virtual machines,” in 2025 IEEE Symposium on Security and Privacy (SP) . IEEE, 2025, pp. 4192–4209
2025
Closest in time.
Google, “gvisor: The container security platform,” 2025, Accessed: 2025-11-10. [Online]. Available: https://gvisor.dev/
2025
Closest in time.
J. Yagnik, “Private ai compute: our next step in building private and helpful ai,” 2025, Accessed: 2025-11-12. [Online]. Available: https://blog.google/technology/ai/google-private-ai-compute/
2025
Closest in time.
Google, “Google private ai compute: Extending on-device privacy with the power of the cloud,” 2025, Accessed: 2025-11-12. [Online]. Available: https://services.google.com/fh/files/misc/private_ai_compute_technical_brief.pdf
2025
Closest in time.
CertiK, “Smart Contract Audit,” 2025. [Online]. Available: https://www.certik.com/products/smart-contract-audit
2025
Closest in time.
J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, and D. Genkin, “ TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ,” in 2026 IEEE Symposium on Security and Privacy (SP) . Los Alamitos, CA, USA: IEEE Computer Society, May 2026, pp. 1894–1912. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/SP63933.2026.00101
2026
Closest in time.