Fetching the paper…
Reading the bibliography…
Large language models (LLMs) possess extensive knowledge and question-answering capabilities, having been widely deployed in privacy-sensitive domains like finance and medical consultation.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Stealing neural networks via timing side channels
Vasisht Duddu, Debasis Samanta, D Vijay Rao, and Valentina E Balas · 2018
Earlier work this paper cites.
Security analysis of deep neural networks operating in the presence of cache side-channel attacks
Sanghyun Hong, Michael Davinroy, Yiǧitcan Kaya, Stuart Nevans Locke, Ian Rackow, Kevin Kulda, Dana Dachman-Soled, and Tudor Dumitraş · 2018
Earlier work this paper cites.
Reverse engineering convolutional neural networks through side-channel information leaks
Weizhe Hua, Zhiru Zhang, and G Edward Suh · 2018
Earlier work this paper cites.
Rendered insecure: Gpu side channel attacks are practical
Hoda Naghibijouybari, Ajaya Neupane, Zhiyun Qian, and Nael Abu-Ghazaleh · 2018
Earlier work this paper cites.
I know what you see: Power side-channel attack on convolutional neural network accelerators
Lingxiao Wei, Bo Luo, Yu Li, Yannan Liu, and Qiang Xu · 2018
Earlier work this paper cites.
Floating-point multiplication timing attack on deep neural network
Gaofeng Dong, Ping Wang, Ping Chen, Ruizhe Gu, and Honggang Hu · 2019
Earlier work this paper cites.
Auditing data provenance in text-generation models
Congzheng Song and Vitaly Shmatikov · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2020
Earlier work this paper cites.
Reverse-engineering deep neural networks using floating-point timing side-channels
Cheng Gongye, Yunsi Fei, and Thomas Wahl · 2020
Earlier work this paper cites.
Deepsniffer: A dnn model extraction framework based on learning architectural hints
Xing Hu, Ling Liang, Shuangchen Li, Lei Deng, Pengfei Zuo, Yu Ji, Xinfeng Xie, Yufei Ding, Chang Liu, Timothy Sherwood, et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Information leakage in embedding models
Congzheng Song and Ananth Raghunathan · 2020
Earlier work this paper cites.
Cache telepathy: Leveraging shared resource attacks to learn { \{ DNN } \} architectures
Mengjia Yan, Christopher W Fletcher, and Josep Torrellas · 2020
Earlier work this paper cites.
Deepem: Deep neural networks model recovery through em side-channel information leakage
Honggang Yu, Haocheng Ma, Kaichen Yang, Yiqiang Zhao, and Yier Jin · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Does bert pretrained on clinical notes reveal sensitive data?
Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron C Wallace · 2021
Earlier work this paper cites.
Leaky nets: Recovering embedded neural network models and inputs through simple power and timing side-channels—attacks and defenses
Saurav Maji, Utsav Banerjee, and Anantha P Chandrakasan · 2021
Earlier work this paper cites.
Invisible probe: Timing attacks with pcie congestion side-channel
Mingtian Tan, Junpeng Wan, Zhe Zhou, and Zhou Li · 2021
Earlier work this paper cites.
Remote power attacks on the versatile tensor accelerator in multi-tenant fpgas
Shanquan Tian, Shayan Moini, Adam Wolnikowski, Daniel Holcomb, Russell Tessier, and Jakub Szefer · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr · 2022
Earlier work this paper cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Mikhail Bosma, Adam Roberts, Maitreya Subbiah, Quoc V Le, and Ilya Sutskever · 2022
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Piloting copilot and codex: Hot temperature, cold prompts, or black magic?
Jean-Baptiste Döderlein, Mathieu Acher, Djamel Eddine Khelladi, and Benoit Combemale · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Earlier work this paper cites.
Ppt: Pre-trained prompt tuning for few-shot learning
Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang · 2022
Earlier work this paper cites.
Are large pre-trained language models leaking your personal information?
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang · 2022
Earlier work this paper cites.
Memorization in nlp fine-tuning methods
Fatemehsadat Mireshghallah, Archit Uniyal, Tianhao Wang, David Evans, and Taylor Berg-Kirkpatrick · 2022
Earlier work this paper cites.
An empirical analysis of memorization in fine-tuned autoregressive language models
Fatemehsadat Mireshghallah, Archit Uniyal, Tianhao Wang, David K Evans, and Taylor Berg-Kirkpatrick · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
Stealthy inference attack on dnn via cache-based side-channel attacks
Han Wang, Syed Mahbub Hafiz, Kartik Patwari, Chen-Nee Chuah, Zubair Shafiq, and Houman Homayoun · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Mona G Flores, Ying Zhang, et al · 2022
Earlier work this paper cites.
Automatic chain of thought prompting in large language models
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola · 2022
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Earlier work this paper cites.
Gptcache: An open-source semantic cache for llm applications enabling faster answers and cost savings
Fu Bang · 2023
Earlier work this paper cites.
A desynchronization-based countermeasure against side-channel analysis of neural networks
Jakub Breier, Dirmanto Jap, Xiaolu Hou, and Shivam Bhasin · 2023
Earlier work this paper cites.
Chatlaw: Open-source legal large language model with integrated external knowledge bases
Jiaxi Cui, Zongjian Li, Yang Yan, Bohua Chen, and Li Yuan · 2023
Earlier work this paper cites.
How continuous batching enables 23x throughput in llm inference while reducing p50 latency, 2023
Cade Daniel, Chen Shen, Eric Liang, and Richard Liaw · 2023
Earlier work this paper cites.
Spy in the gpu-box: Covert and side channel attacks on multi-gpu systems
Sankha Baran Dutta, Hoda Naghibijouybari, Arjun Gupta, Nael Abu-Ghazaleh, Andres Marquez, and Kevin Barker · 2023
Earlier work this paper cites.
Comave: Contrastive pre-training with multi-scale masking for attribute value extraction
Xinnan Guo, Wentao Deng, Yongrui Chen, Yang Li, Mengdi Zhou, Guilin Qi, Tianxing Wu, Dong Yang, Liubin Wang, and Yong Pan · 2023
Earlier work this paper cites.
Lm-infinite: Simple on-the-fly length generalization for large language models
Chi Han, Qifan Wang, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang · 2023
Earlier work this paper cites.
Peter Horvath, Lukasz Chmielewski, Leo Weissbart, Lejla Batina, and Yuval Yarom · 2023
Earlier work this paper cites.
How to better configure your cache; GPTCache
Zilliz Inc · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Sentence embedding leaks more information than you expect: Generative embedding inversion attack to recover the whole sentence
Haoran Li, Mingshi Xu, and Yangqiu Song · 2023
Cited alongside, same era.
Structured chain-of-thought prompting for code generation
Jia Li, Ge Li, Yongmin Li, and Zhi Jin · 2023
Cited alongside, same era.
Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge
Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang · 2023
Cited alongside, same era.
Privacy-aware semantic cache for large language models
Waris Gill, Mohamed Elidrisi, Pallavi Kalapatapu, Ali Anwar, and Muhammad Ali Gulzar · 2024
Closest in time.
Prompt cache: Modular attention reuse for low-latency inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong · 2024
Closest in time.
Gemini API : Google AI for Developers
Google · 2024
Closest in time.
Hydragen: High-throughput llm inference with shared prefixes
Jordan Juravsky, Bradley Brown, Ryan Ehrlich, Daniel Y Fu, Christopher Ré, and Azalia Mirhoseini · 2024
Closest in time.
Rishi Kalra, Zekun Wu, Ayesha Gulley, Airlie Hilliard, Xin Guan, Adriano Koshiyama, and Philip Treleaven · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nils Lukas, Ahmed Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-Béguelin · 2023
Cited alongside, same era.
Prompt engineering in large language models
Ggaliwango Marvin, Nakayiza Hellen, Daudi Jjingo, and Joyce Nakatumba-Nabende · 2023
Cited alongside, same era.
Membership inference attacks against language models via neighbourhood comparison
Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schoelkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick · 2023
Cited alongside, same era.
Large language models challenge the future of higher education
Silvia Milano, Joshua A McGrane, and Sabina Leonelli · 2023
Cited alongside, same era.
John X Morris, Wenting Zhao, Justin T Chiu, Vitaly Shmatikov, and Alexander M Rush · 2023
Cited alongside, same era.
Text embeddings reveal (almost) as much as text
John Xavier Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M Rush · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean · 2023
Cited alongside, same era.
Closest in time.
Improve speed and reduce cost for generative AI workloads with a persistent semantic cache in Amazon MemoryDB
Santiago Flores Kanter · 2024
Closest in time.
Propile: Probing privacy leakage in large language models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh · 2024
Closest in time.
Scalm: Towards semantic caching for automated chat services with large language models
Jiaxing Li, Chi Xu, Feng Wang, Isaac M von Riedemann, Cong Zhang, and Jiangchuan Liu · 2024
Closest in time.
Why are my prompts leaked? unraveling prompt extraction threats in customized large language models
Zi Liang, Haibo Hu, Qingqing Ye, Yaxin Xiao, and Haoyang Li · 2024
Closest in time.
Student interaction with newtbot: An llm-as-tutor chatbot for secondary physics education
Anna Lieb and Toshali Goel · 2024
Closest in time.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han · 2024
Closest in time.
Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models
Shweta Mahajan, Tanzila Rahman, Kwang Moo Yi, and Leonid Sigal · 2024
Closest in time.
Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models
Shweta Mahajan, Tanzila Rahman, Kwang Moo Yi, and Leonid Sigal · 2024
Closest in time.
Did the neurons read your book? document-level membership inference for large language models
Matthieu Meeus, Shubham Jain, Marek Rei, and Yves-Alexandre de Montjoye · 2024
Closest in time.
Context-based semantic caching for llm applications
Ramaswami Mohandoss · 2024
Closest in time.
Using the Context Caching Feature of the Kimi API
MoonShot · 2024
Closest in time.
Sfr-rag: Towards contextually faithful llms
Xuan-Phi Nguyen, Shrey Pandit, Senthil Purushwalkam, Austin Xu, Hailin Chen, Yifei Ming, Zixuan Ke, Silvio Savarese, Caiming Xong, and Shafiq Joty · 2024
Closest in time.
Ragged Batching; NVIDIA Triton Inference Server
NVIDIA · 2024
Closest in time.
Introducing the GPT Store We’re launching the GPT Store to help you find
OpenAI · 2024
Closest in time.
Learning to Reason with LLMs We are introducing OpenAI o1, a new large language model trained with LLMs
OpenAI · 2024
Closest in time.
Prompt caching: Reduce latency and cost with prompt caching
OpenAI · 2024
Closest in time.
Cache (Simple & Semantic) - Portkey Docs
Portkey · 2024
Closest in time.
A systematic survey of prompt engineering in large language models: Techniques and applications
Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Vinija Jain, Samrat Mondal, and Aman Chadha · 2024
Closest in time.
Prompt stealing attacks against large language models
Zeyang Sha and Yang Zhang · 2024
Closest in time.
Implementing Semantic Caching: A Step-by-Step Guide to Faster, Cost-Effective GenAI Workflows
Arun Shankar · 2024
Closest in time.
Prompt stealing attacks against { \{ Text-to-Image } \} generation models
Xinyue Shen, Yiting Qu, Michael Backes, and Yang Zhang · 2024
Closest in time.
Legal-lm: Knowledge graph enhanced large language models for law consulting
Juanming Shi, Qinglang Guo, Yong Liao, Yuxing Wang, Shijia Chen, and Shenglin Liang · 2024
Closest in time.
Optimize Azure OpenAI Applications with Semantic Caching
Sudarsan · 2024
Closest in time.
Lawluo: A chinese law firm co-run by llm agents
Jingyun Sun, Chengxiao Dai, Zhongze Luo, Yangbo Chang, and Yang Li · 2024
Closest in time.
GPT for Microsoft Excel or Google Sheets
Talarian · 2024
Closest in time.
OpenAI GPT prompt generator
Talarian · 2024
Closest in time.
Scaling laws with vocabulary: Larger models deserve larger vocabularies
Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong · 2024
Closest in time.
Zuoshouyisheng open platforms
Zuoyi Technology · 2024
Closest in time.
Towards generalist biomedical ai
Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Charles Lau, Ryutaro Tanno, Ira Ktena, et al · 2024
Closest in time.
What was your prompt? a remote keylogging attack on ai assistants
Roy Weiss, Daniel Ayzenshteyn, Guy Amit, and Yisroel Mirsky · 2024
Closest in time.
What was your prompt? a remote keylogging attack on AI assistants
Roy Weiss, Daniel Ayzenshteyn, and Yisroel Mirsky · 2024
Closest in time.
Ai for education (ai4edu): Advancing personalized education with llm and adaptive learning
Qingsong Wen, Jing Liang, Carles Sierra, Rose Luckin, Richard Tong, Zitao Liu, Peng Cui, and Jiliang Tang · 2024
Closest in time.
Yang Wu, Chenghao Wang, Ece Gumusel, and Xiaozhong Liu · 2024
Closest in time.
Rule: Reliable multimodal rag for factuality in medical vision language models
Peng Xia, Kangyu Zhu, Haoran Li, Hongtu Zhu, Yun Li, Gang Li, Linjun Zhang, and Huaxiu Yao · 2024
Closest in time.
Prsa: Prompt reverse stealing attacks against large language models
Yong Yang, Xuhong Zhang, Yi Jiang, Xi Chen, Haoyu Wang, Shouling Ji, and Zonghui Wang · 2024
Closest in time.
Chunkattention: Efficient self-attention with prefix-aware kv cache and two-phase partition
Lu Ye, Ze Tao, Yong Huang, and Yang Li · 2024
Closest in time.
Wkvquant: Quantizing weight and key/value cache for large language models gains more
Yuxuan Yue, Zhihang Yuan, Haojie Duanmu, Sifan Zhou, Jianlong Wu, and Liqiang Nie · 2024
Closest in time.
Subgen: Token generation in sublinear time and memory
Amir Zandieh, Insu Han, Vahab Mirrokni, and Amin Karbasi · 2024
Closest in time.
Effective prompt extraction from language models
Yiming Zhang, Nicholas Carlini, and Daphne Ippolito · 2024
Closest in time.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2024
Closest in time.
Efficient prompt caching via embedding similarity, 2024
Hanlin Zhu, Banghua Zhu, and Jiantao Jiao · 2024
Closest in time.
Henry Peng Zou, Vinay Samuel, Yue Zhou, Weizhi Zhang, Liancheng Fang, Zihe Song, Philip S Yu, and Cornelia Caragea · 2024
Closest in time.