Fetching the paper…
Reading the bibliography…
With the rapid development of artificial intelligence, large language models (LLMs) have made remarkable advancements in natural language processing.
Medical document anonymization with a semantic lexicon.. In Proceedings of the AMIA Symposium . American Medical Informatics Association, 729
Patrick Ruch, Robert H Baud, Anne-Marie Rassinoux, Pierrette Bouillon, and Gilbert Robert. 2000 · 2000
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision . 19–27
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security . 308–318
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016 · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data. In Artificial intelligence and statistics . PMLR, 1273–1282
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017 · 2017
Earlier work this paper cites.
Fine-pruning: Defending against backdooring attacks on deep neural networks. In International symposium on research in attacks, intrusions, and defenses . Springer, 273–294
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2018 · 2018
Earlier work this paper cites.
Exploiting unintended feature leakage in collaborative learning. In 2019 IEEE symposium on security and privacy (SP) . IEEE, 691–706
Luca Melis, Congzheng Song, Emiliano De Cristofaro, and Vitaly Shmatikov. 2019 · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 2019
Earlier work this paper cites.
Weight Poisoning Attacks on Pretrained Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 2793–2806
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks. In International Conference on Learning Representations
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2020 · 2020
Earlier work this paper cites.
Improving the accuracy of medical diagnosis with causal machine learning
Jonathan G Richens, Ciarán M Lee, and Saurabh Johri. 2020 · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 4222–4235
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
{ \{ T-Miner } \} : A generative approach to defend against trojan attacks on { \{ DNN-based } \} text classification. In 30th USENIX Security Symposium (USENIX Security 21) . 2255–2272
Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K Reddy, and Bimal Viswanath. 2021 · 2021
Earlier work this paper cites.
Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21) . 2633–2650
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
How should pre-trained language models be fine-tuned towards adversarial robustness?
Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang. 2021 · 2021
Earlier work this paper cites.
Design and evaluation of a multi-domain trojan detection method on deep neural networks
Yansong Gao, Yeonjae Kim, Bao Gia Doan, Zhi Zhang, Gongxuan Zhang, Surya Nepal, Damith C Ranasinghe, and Hyoungshick Kim. 2021 · 2021
Earlier work this paper cites.
Gradient-based Adversarial Attacks against Text Transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 5747–5757
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. 2021 · 2021
Earlier work this paper cites.
Membership inference attack susceptibility of clinical language models
Abhyuday Jagannatha, Bhanu Pratap Singh Rawat, and Hong Yu. 2021 · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 3045–3059
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
Large Language Models Can Be Strong Differentially Private Learners. In International Conference on Learning Representations
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2021 · 2021
Earlier work this paper cites.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021 · 2021
Earlier work this paper cites.
Probing toxic content in large pre-trained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 4262–4274
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit-Yan Yeung. 2021 · 2021
Earlier work this paper cites.
ONION: A Simple and Effective Defense Against Textual Backdoor Attacks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 9558–9566
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2021a · 2021
Earlier work this paper cites.
Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 4569–4580
Fanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li, Zhiyuan Liu, and Maosong Sun. 2021b · 2021
Earlier work this paper cites.
Bddr: An effective defense against textual backdoor attacks
Kun Shao, Junan Yang, Yang Ai, Hui Liu, and Yu Zhang. 2021 · 2021
Earlier work this paper cites.
Concealed Data Poisoning Attacks on NLP Models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. 2021 · 2021
Earlier work this paper cites.
Challenges in Detoxifying Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2021 . 2447–2469
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Lamp: Extracting text from gradients with language model priors
Mislav Balunovic, Dimitar Dimitrov, Nikola Jovanović, and Martin Vechev. 2022 · 2022
Earlier work this paper cites.
Badprompt: Backdoor attacks on continuous prompts
Xiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang, et al · 2022
Earlier work this paper cites.
Quantifying Memorization Across Neural Language Models. In The Eleventh International Conference on Learning Representations
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Earlier work this paper cites.
THE-X: Privacy-Preserving Transformer Inference with Homomorphic Encryption. In Findings of the Association for Computational Linguistics: ACL 2022 . 3510–3520
Tianyu Chen, Hangbo Bao, Shaohan Huang, Li Dong, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei. 2022 · 2022
Earlier work this paper cites.
A unified evaluation of textual backdoor learning: Frameworks and benchmarks
Ganqu Cui, Lifan Yuan, Bingxiang He, Yangyi Chen, Zhiyuan Liu, and Maosong Sun. 2022 · 2022
Earlier work this paper cites.
Perd: Perturbation sensitivity-based neural trojan detection framework on nlp applications
Diego Garcia-soto, Huili Chen, and Farinaz Koushanfar. 2022 · 2022
Earlier work this paper cites.
Protecting intellectual property of language generation apis with lexical watermark. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 10758–10766
Xuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu, and Chenguang Wang. 2022 · 2022
Earlier work this paper cites.
Scaling laws and interpretability of learning from repeated data
Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al · 2022
Earlier work this paper cites.
Deduplicating training data mitigates privacy risks in language models. In International Conference on Machine Learning . PMLR, 10697–10707
Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022 · 2022
Earlier work this paper cites.
You Don’t Know My Favorite Color: Preventing Dialogue Representations from Revealing Speakers’ Private Personas. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 5858–5870
Haoran Li, Yangqiu Song, and Lixin Fan. 2022a · 2022
Earlier work this paper cites.
P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . Association for Computational Linguistics
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022 · 2022
Earlier work this paper cites.
Paradetox: Detoxification with parallel data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 6804–6818
Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022 · 2022
Earlier work this paper cites.
Differentially private decoding in large language models
Jimit Majmudar, Christophe Dupuy, Charith Peris, Sami Smaili, Rahul Gupta, and Richard Zemel. 2022 · 2022
Earlier work this paper cites.
Differentially Private Language Models for Secure Data Sharing. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 4860–4873
Justus Mattern, Zhijing Jin, Benjamin Weggenmann, Bernhard Schoelkopf, and Mrinmaya Sachan. 2022 · 2022
Earlier work this paper cites.
Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 8332–8347
Fatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, and Reza Shokri. 2022 · 2022
Earlier work this paper cites.
Exploring Cross-lingual Text Detoxification with Large Multilingual Language Models.. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop . 346–354
Daniil Moskovskiy, Daryna Dementieva, and Alexander Panchenko. 2022 · 2022
Earlier work this paper cites.
Canary extraction in natural language understanding models
Rahil Parikh, Christophe Dupuy, and Rahul Gupta. 2022 · 2022
Earlier work this paper cites.
Ignore Previous Prompt: Attack Techniques For Language Models. In NeurIPS ML Safety Workshop
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
Defending against stealthy backdoor attacks
Sangeet Sagar, Abhinav Bhatt, and Abhijith Srinivas Bidaralli. 2022 · 2022
Earlier work this paper cites.
Rethink stealthy backdoor attacks in natural language processing
Lingfeng Shen, Haiyun Jiang, Lemao Liu, and Shuming Shi. 2022a · 2022
Earlier work this paper cites.
A survey on backdoor attack and defense in natural language processing. In 2022 IEEE 22nd International Conference on Software Quality, Reliability and Security (QRS) . IEEE, 809–820
Xuan Sheng, Zhaoyang Han, Piji Li, and Xiangmao Chang. 2022 · 2022
Earlier work this paper cites.
Just Fine-tune Twice: Selective Differential Privacy for Large Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 6327–6340
Weiyan Shi, Ryan Shea, Si Chen, Chiyuan Zhang, Ruoxi Jia, and Zhou Yu. 2022 · 2022
Earlier work this paper cites.
Seqpate: Differentially private text generation via knowledge distillation
Zhiliang Tian, Yingxiu Zhao, Ziyue Huang, Yu-Xiang Wang, Nevin L Zhang, and He He. 2022 · 2022
Earlier work this paper cites.
Emergent Abilities of Large Language Models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Earlier work this paper cites.
Enhanced membership inference attacks against machine learning models. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security . 3093–3106
Jiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler, and Reza Shokri. 2022 · 2022
Earlier work this paper cites.
Adversarial attacks and defenses in deep learning: From a perspective of cybersecurity
Shuai Zhou, Chi Liu, Dayong Ye, Tianqing Zhu, Wanlei Zhou, and Philip S. Yu. 2022 · 2022
Earlier work this paper cites.
Label-only model inversion attacks: Attack with the least information
Tianqing Zhu, Dayong Ye, Shuai Zhou, Bo Liu, and Wanlei Zhou. 2022 · 2022
Earlier work this paper cites.
Tmi! finetuned models leak private information from their pretraining data
John Abascal, Stanley Wu, Alina Oprea, and Jonathan Ullman. 2023 · 2023
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions. In The Twelfth International Conference on Learning Representations
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Rottger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. 2023 · 2023
Cited alongside, same era.
Are aligned neural networks adversarially aligned?. In 37th Annual Conference on Neural Information Processing Systems (NeurIPS 2023)
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramèr, and Ludwig Schmidt. 2023 · 2023
Cited alongside, same era.
Jailbreaker in jail: Moving target defense for large language models. In Proceedings of the 10th ACM Workshop on Moving Target Defense . 29–32
Bocheng Chen, Advait Paliwal, and Qiben Yan. 2023 · 2023
Cited alongside, same era.
ChatGPT: A Threat Against the CIA Triad of Cyber Security. In 2023 IEEE International Conference on Electro Information Technology (eIT) . IEEE, 1–6
MD Minhaz Chowdhury, Nafiz Rifat, Mostofa Ahsan, Shadman Latif, Rahul Gomes, and Md Saifur Rahman. 2023 · 2023
Cited alongside, same era.
Evil geniuses: Delving into the safety of llm-based agents
Yu Tian, Xiao Yang, Jingyuan Zhang, Yinpeng Dong, and Hang Su. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Poisoning language models during instruction tuning. In International Conference on Machine Learning . PMLR, 35413–35425
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023 · 2023
Later among the works it cites.
Adversarial demonstration attacks on large language models
Jiongxiao Wang, Zichen Liu, Keun Hee Park, Muhao Chen, and Chaowei Xiao. 2023b · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2023 · 2023
Cited alongside, same era.
Yimo Deng and Huangxun Chen. 2023 · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 1236–1270
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023 · 2023
Cited alongside, same era.
Puma: Secure inference of llama-7b in five minutes
Ye Dong, Wen-jie Lu, Yancheng Zheng, Haoqi Wu, Derun Zhao, Jin Tan, Zhicong Huang, Cheng Hong, Tao Wei, and Wenguang Cheng. 2023 · 2023
Cited alongside, same era.
Uor: Universal backdoor attacks on pre-trained language models
Wei Du, Peixuan Li, Boqun Li, Haodong Zhao, and Gongshen Liu. 2023 · 2023
Cited alongside, same era.
Who’s Harry Potter? Approximate Unlearning in LLMs
Ronen Eldan and Mark Russinovich. 2023 · 2023
Cited alongside, same era.
Fate-llm: A industrial grade federated learning framework for large language models
Tao Fan, Yan Kang, Guoqiang Ma, Weijing Chen, Wenbin Wei, Lixin Fan, and Qiang Yang. 2023 · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security . 79–90
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Cited alongside, same era.
Jiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik, and Chaowei Xiao. 2023c · 2023
Later among the works it cites.
CASSOCK: Viable backdoor attacks against DNN in the wall of source-specific backdoor defenses. In Proceedings of the 2023 ACM Asia Conference on Computer and Communications Security . 938–950
Shang Wang, Yansong Gao, Anmin Fu, Zhi Zhang, Yuqing Zhang, Willy Susilo, and Dongxi Liu. 2023a · 2023
Later among the works it cites.
Lmsanitator: Defending prompt-tuning against task-agnostic backdoors
Chengkun Wei, Wenlong Meng, Zhikun Zhang, Min Chen, Minghu Zhao, Wenjing Fang, Lei Wang, Zihui Zhang, and Wenzhi Chen. 2023a · 2023
Later among the works it cites.
Jailbreak and guard aligned language models with only few in-context demonstrations
Zeming Wei, Yifei Wang, and Yisen Wang. 2023b · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua. 2023 · 2023
Later among the works it cites.
Unveiling security, privacy, and ethical concerns of chatgpt
Xiaodong Wu, Ran Duan, and Jianbing Ni. 2023 · 2023
Later among the works it cites.
Large language models can be good privacy protection learners
Yijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu, Xianjun Yang, Xiao Luo, Wenchao Yu, Xujiang Zhao, Yanchi Liu, Haifeng Chen, et al · 2023
Later among the works it cites.
Training large-vocabulary neural language models by private federated learning for resource-constrained devices. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Mingbin Xu, Congzheng Song, Ye Tian, Neha Agrawal, Filip Granqvist, Rogier van Dalen, Xiao Zhang, Arturo Argueta, Shiyi Han, Yaqiao Deng, et al · 2023
Later among the works it cites.
BITE: Textual Backdoor Attacks with Iterative Trigger Injection. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 12951–12968
Jun Yan, Vansh Gupta, and Xiang Ren. 2023 · 2023
Later among the works it cites.
Local differential privacy and its applications: A comprehensive survey
Mengmeng Yang, Taolin Guo, Tianqing Zhu, Ivan Tjuawinata, Jun Zhao, and Kwok-Yan Lam. 2023 · 2023
Later among the works it cites.
Large Language Model Unlearning. In Socially Responsible Language Modelling Research
Yuanshun Yao, Xiaojun Xu, and Yang Liu. 2023 · 2023
Later among the works it cites.
Selective pre-training for private fine-tuning
Da Yu, Sivakanth Gopi, Janardhan Kulkarni, Zinan Lin, Saurabh Naik, Tomasz Lukasz Religa, Jian Yin, and Huishuai Zhang. 2023a · 2023
Later among the works it cites.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2023 · 2023
Later among the works it cites.
Remark-llm: A robust and efficient watermarking framework for generative large language models
Ruisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, and Farinaz Koushanfar. 2023a · 2023
Later among the works it cites.
Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success
Yiming Zhang and Daphne Ippolito. 2023 · 2023
Later among the works it cites.
Zhexin Zhang, Jiaxin Wen, and Minlie Huang. 2023b · 2023
Later among the works it cites.
Fedpetuning: When federated learning meets the parameter-efficient tuning methods of pre-trained language models. In Annual Meeting of the Association of Computational Linguistics 2023 . Association for Computational Linguistics (ACL), 9963–9977
Zhuo Zhang, Yuanhang Yang, Yong Dai, Qifan Wang, Yue Yu, Lizhen Qu, and Zenglin Xu. 2023c · 2023
Later among the works it cites.
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 12303–12317
Shuai Zhao, Jinming Wen, Anh Luu, Junbo Zhao, and Jie Fu. 2023b · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
Sara Abdali, Richard Anarfi, CJ Barberan, and Jia He. 2024 · 2024
Closest in time.
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
Tim Baumgärtner, Yang Gao, Dana Alon, and Donald Metzler. 2024 · 2024
Closest in time.
Sukmin Cho, Soyeong Jeong, Jeongyeon Seo, Taeho Hwang, and Jong C Park. 2024 · 2024
Closest in time.
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models
Arijit Ghosh Chowdhury, Md Mofijul Islam, Vaibhav Kumar, Faysal Hossain Shezan, Vinija Jain, and Aman Chadha. 2024 · 2024
Closest in time.
Risk taxonomy, mitigation, and assessment benchmarks of large language model systems
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al · 2024
Closest in time.
Security and privacy challenges of large language models: A survey
Badhan Chandra Das, M Hadi Amini, and Yanzhao Wu. 2024 · 2024
Closest in time.
Safeguarding Large Language Models: A Survey
Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, et al · 2024
Closest in time.
Attacks, defenses and evaluations for llm conversation safety: A survey
Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, and Yu Qiao. 2024b · 2024
Closest in time.
Flocks of stochastic parrots: Differentially private prompt learning for large language models
Haonan Duan, Adam Dziedzic, Nicolas Papernot, and Franziska Boenisch. 2024 · 2024
Closest in time.
Privacy Preserving Prompt Engineering: A Survey
Kennedy Edemacu and Xintao Wu. 2024 · 2024
Closest in time.
On the Effectiveness of Distillation in Mitigating Backdoors in Pre-trained Encoder
Tingxu Han, Shenghan Huang, Ziqi Ding, Weisong Sun, Yebo Feng, Chunrong Fang, Jun Li, Hanwei Qian, Cong Wu, Quanjun Zhang, et al · 2024
Closest in time.
Propile: Probing privacy leakage in large language models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2024 · 2024
Closest in time.
Double-I Watermark: Protecting Model Copyright for LLM Fine-tuning
Shen Li, Liuyi Yao, Jinyang Gao, Lan Zhang, and Yaliang Li. 2024c · 2024
Closest in time.
Badedit: Backdooring large language models by model editing
Yanzhou Li, Tianlin Li, Kangjie Chen, Jian Zhang, Shangqing Liu, Wenhan Wang, Tianwei Zhang, and Yang Liu. 2024a · 2024
Closest in time.
Personal llm agents: Insights and survey about the capability, efficiency and security
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al · 2024
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2024 · 2024
Closest in time.
Machine unlearning for generative AI
Yashaswini Viswanath, Sudha Jamthe, Suresh Lokiah, and Emanuele Bianchini. 2024 · 2024
Closest in time.
Large language models for cyber security: A systematic literature review
HanXiang Xu, ShenAo Wang, Ningke Li, Yanjie Zhao, Kai Chen, Kailong Wang, Yang Liu, Ting Yu, and HaoYu Wang. 2024b · 2024
Closest in time.
On Protecting the Data Privacy of Large Language Models (LLMs): A Survey
Biwei Yan, Kun Li, Minghui Xu, Yueyan Dong, Yue Zhang, Zhaochun Ren, and Xiuzheng Cheng. 2024 · 2024
Closest in time.
A comprehensive overview of backdoor attacks in large language models within communication networks
Haomiao Yang, Kunlan Xiang, Mengyu Ge, Hongwei Li, Rongxing Lu, and Shui Yu. 2024c · 2024
Closest in time.
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun. 2024a · 2024
Closest in time.
Sneakyprompt: Jailbreaking text-to-image generative models. In 2024 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, 123–123
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao. 2024b · 2024
Closest in time.
Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4485–4489
Dongyu Yao, Jianshu Zhang, Ian G Harris, and Marcel Carlsson. 2024c · 2024
Closest in time.
Poisonprompt: Backdoor attack on prompt-based large language models. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7745–7749
Hongwei Yao, Jian Lou, and Zhan Qin. 2024b · 2024
Closest in time.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024a · 2024
Closest in time.
Defending against Label-only Attacks via Meta-Reinforcement Learning
Dayong Ye, Tianqing Zhu, Kun Gao, and Wanlei Zhou. 2024 · 2024
Closest in time.
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
Shenglai Zeng, Jiankun Zhang, Pengfei He, Yue Xing, Yiding Liu, Han Xu, Jie Ren, Shuaiqiang Wang, Dawei Yin, Yi Chang, et al · 2024
Closest in time.
Human-Imperceptible Retrieval Poisoning Attacks in LLM-Powered Applications
Quan Zhang, Binqi Zeng, Chijin Zhou, Gwihwan Go, Heyuan Shi, and Yu Jiang. 2024 · 2024
Closest in time.
Inversion-guided Defense: Detecting Model Stealing Attacks by Output Inverting
Shuai Zhou, Tianqing Zhu, Dayong Ye, Wanlei Zhou, and Wei Zhao. 2024 · 2024
Closest in time.
PoisonedRAG: Knowledge Poisoning Attacks to Retrieval-Augmented Generation of Large Language Models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2024 · 2024
Closest in time.
Robust multi-bit natural language watermarking through invariant features. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2092–2115
KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak. 2023 · 2092
Closest in time.