Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated extraordinary capabilities and contributed to multiple fields, such as generating and summarizing text, language translation, and question-answering.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
Roger Ratcliff. 1990 · 1990
Earlier work this paper cites.
Health insurance portability and accountability act (HIPPA) compliant access control model for web services
Vivying SY Cheng et al · 2006
Earlier work this paper cites.
Privacy-preserving face recognition. In Privacy Enhancing Technologies: 9th International Symposium, PETS 2009, Seattle, WA, USA, August 5-7, 2009. Proceedings 9 . Springer, 235–253
Zekeriya Erkin, Martin Franz, Jorge Guajardo, Stefan Katzenbeisser, Inald Lagendijk, and Tomas Toft. 2009 · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing . 1631–1642
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Exposing mobile malware from the inside (or what is your mobile app really doing?)
Dimitrios Damopoulos, Georgios Kambourakis, Stefanos Gritzalis, and Sang Oh Park. 2014 · 2014
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Earlier work this paper cites.
Secure multiparty computation
Ronald Cramer, Ivan Bjerre Damgård, et al · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2015 · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security . 308–318
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016 · 2016
Earlier work this paper cites.
A literature survey on social engineering attacks: Phishing attack. In 2016 international conference on computing, communication and automation (ICCCA) . IEEE, 537–540
Surbhi Gupta, Abhishek Singhal, and Akanksha Kapoor. 2016 · 2016
Earlier work this paper cites.
Neural architectures for named entity recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016 · 2016
Earlier work this paper cites.
ReCon: Revealing and controlling PII leaks in mobile network traffic. In Proceedings of the 14th Annual International Conference on Mobile Systems, Applications, and Services . 361–374
Jingjing Ren, Ashwin Rao, Martina Lindorfer, Arnaud Legout, and David Choffnes. 2016 · 2016
Earlier work this paper cites.
"Why should I trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2818–2826
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al · 2016
Earlier work this paper cites.
Privacy-preserving deep learning via additively homomorphic encryption
Yoshinori Aono, Takuya Hayashi, Lihua Wang, Shiho Moriai, et al · 2017
Earlier work this paper cites.
Mitigating poisoning attacks on machine learning models: A data provenance based approach. In Proceedings of the 10th ACM workshop on artificial intelligence and security . 103–110
Nathalie Baracaldo, Bryant Chen, Heiko Ludwig, and Jaehoon Amir Safavi. 2017 · 2017
Earlier work this paper cites.
Obfuscation-Resilient Privacy Leak Detection for Mobile Apps Through Differential Analysis.. In NDSS , Vol. 17. 10–14722
Andrea Continella, Yanick Fratantonio, Martina Lindorfer, Alessandro Puccetti, Ali Zand, Christopher Kruegel, Giovanni Vigna, et al · 2017
Earlier work this paper cites.
Differentially private federated learning: A client level perspective
Robin C Geyer, Tassilo Klein, and Moin Nabi. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott Lundberg. 2017 · 2017
Earlier work this paper cites.
HDBSCAN: Hierarchical density based clustering
Leland McInnes, John Healy, and Steve Astels. 2017 · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision . 618–626
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP) . IEEE, 3–18
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
The eu general data protection regulation (GDPR)
Paul Voigt and Axel Von dem Bussche. 2017 · 2017
Earlier work this paper cites.
Detecting backdoor attacks on deep neural networks by activation clustering
Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Benjamin Edwards, Taesung Lee, Ian Molloy, and Biplav Srivastava. 2018 · 2018
Earlier work this paper cites.
SentiNet: Detecting physical attacks against deep learning systems.(2018)
Edward Chou, Florian Tramèr, Giancarlo Pellegrino, and Dan Boneh. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton. 2018 · 2018
Earlier work this paper cites.
Trojaning attack on neural networks. In 25th Annual Network And Distributed System Security Symposium (NDSS 2018) . Internet Soc
Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang. 2018b · 2018
Earlier work this paper cites.
Machine learning with membership privacy using adversarial regularization. In Proceedings of the 2018 ACM SIGSAC conference on computer and communications security . 634–646
Milad Nasr, Reza Shokri, and Amir Houmansadr. 2018 · 2018
Earlier work this paper cites.
Ahmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang, Mario Fritz, and Michael Backes. 2018 · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting. In 2018 IEEE 31st computer security foundations symposium (CSF) . IEEE, 268–282
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. 2018 · 2018
Earlier work this paper cites.
DeepInspect: A Black-box Trojan Detection and Mitigation Framework for Deep Neural Networks.. In IJCAI , Vol. 2. 8
Huili Chen, Cheng Fu, Jishen Zhao, and Farinaz Koushanfar. 2019 · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing. In international conference on machine learning . PMLR, 1310–1320
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter. 2019 · 2019
Earlier work this paper cites.
Strip: A defence against trojan attacks on deep neural networks. In Proceedings of the 35th Annual Computer Security Applications Conference . 113–125
Yansong Gao, Change Xu, Derui Wang, Shiping Chen, Damith C Ranasinghe, and Surya Nepal. 2019 · 2019
Earlier work this paper cites.
Efficient adaptation of pretrained transformers for abstractive summarization
Andrew Hoang, Antoine Bosselut, Asli Celikyilmaz, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Memguard: Defending against black-box membership inference attacks via adversarial examples. In Proceedings of the 2019 ACM SIGSAC conference on computer and communications security . 259–274
Jinyuan Jia, Ahmed Salem, Michael Backes, Yang Zhang, and Neil Zhenqiang Gong. 2019 · 2019
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT , Vol. 1. 2
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
White-box vs black-box: Bayes optimal strategies for membership inference. In International Conference on Machine Learning . PMLR, 5558–5567
Alexandre Sablayrolles, Matthijs Douze, Cordelia Schmid, Yann Ollivier, and Hervé Jégou. 2019 · 2019
Earlier work this paper cites.
Auditing data provenance in text-generation models. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 196–206
Congzheng Song and Vitaly Shmatikov. 2019 · 2019
Earlier work this paper cites.
Identifying and mitigating backdoor attacks in neural networks. In 2019 IEEE Symposium on Security and Privacy (SP) . IEEE, 707–723
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao. 2019a · 2019
Earlier work this paper cites.
A survey of zero-shot learning: Settings, methods, and applications
Wei Wang, Vincent W Zheng, Han Yu, and Chunyan Miao. 2019b · 2019
Earlier work this paper cites.
DIALOGPT: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2019 · 2019
Earlier work this paper cites.
Deep leakage from gradients
Ligeng Zhu, Zhijian Liu, and Song Han. 2019 · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 2019
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Earlier work this paper cites.
GPT-3: Its nature, scope, limits, and consequences
Luciano Floridi and Massimo Chiriatti. 2020 · 2020
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2020 · 2020
Earlier work this paper cites.
Inverting gradients-how easy is it to break privacy in federated learning?
Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, and Michael Moeller. 2020 · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Earlier work this paper cites.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
KART: Privacy leakage framework of language models pre-trained with clinical records
Yuta Nakamura, Shouhei Hanaoka, Yukihiro Nomura, Naoto Hayashi, Osamu Abe, Shuntaro Yada, Shoko Wakamiya, and Eiji Aramaki. 2020 · 2020
Earlier work this paper cites.
Privacy risks of general-purpose language models. In 2020 IEEE Symposium on Security and Privacy (SP) . IEEE, 1314–1331
Xudong Pan, Mi Zhang, Shouling Ji, and Min Yang. 2020 · 2020
Earlier work this paper cites.
ONION: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Adversarial attacks and defenses in deep learning
Kui Ren, Tianhang Zheng, Zhan Qin, and Xue Liu. 2020 · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
Information leakage in embedding models. In Proceedings of the 2020 ACM SIGSAC conference on computer and communications security . 377–390
Congzheng Song and Ananth Raghunathan. 2020 · 2020
Earlier work this paper cites.
Concealed data poisoning attacks on NLP models
Eric Wallace, Tony Z Zhao, Shi Feng, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
Generalizing from a few examples: A survey on few-shot learning
Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni. 2020 · 2020
Earlier work this paper cites.
A framework for evaluating client privacy leakages in federated learning. In Computer Security–ESORICS 2020: 25th European Symposium on Research in Computer Security, ESORICS 2020, Guildford, UK, September 14–18, 2020, Proceedings, Part I 25 . Springer, 545–566
Wenqi Wei, Ling Liu, Margaret Loper, Ka-Ho Chow, Mehmet Emre Gursoy, Stacey Truex, and Yanzhao Wu. 2020 · 2020
Earlier work this paper cites.
Analyzing information leakage of updates to natural language models. In Proceedings of the 2020 ACM SIGSAC conference on computer and communications security . 363–375
Santiago Zanella-Béguelin, Lukas Wutschitz, Shruti Tople, Victor Rühle, Andrew Paverd, Olga Ohrimenko, Boris Köpf, and Marc Brockschmidt. 2020 · 2020
Earlier work this paper cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li. 2020 · 2020
Earlier work this paper cites.
Large-scale differentially private BERT
Rohan Anil, Badih Ghazi, Vineet Gupta, Ravi Kumar, and Pasin Manurangsi. 2021 · 2021
Earlier work this paper cites.
Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21) . 2633–2650
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
TAG: Gradient attack on transformer-based language models
Jieren Deng, Yijue Wang, Ji Li, Chao Shang, Hang Liu, Sanguthevar Rajasekaran, and Caiwen Ding. 2021 · 2021
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Alexandre Sablayrolles Hervé Jégou Guo, Chuan and Douwe Kiela. 2021 · 2021
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. 2021 · 2021
Earlier work this paper cites.
Privacy analysis in language models via training data leakage report
Huseyin A Inan, Osman Ramadan, Lukas Wutschitz, Daniel Jones, Victor Rühle, James Withers, and Robert Sim. 2021a · 2021
Earlier work this paper cites.
Training data leakage analysis in language models
Huseyin A Inan, Osman Ramadan, Lukas Wutschitz, Daniel Jones, Victor Rühle, James Withers, and Robert Sim. 2021b · 2021
Earlier work this paper cites.
Membership inference attack susceptibility of clinical language models
Abhyuday Jagannatha, Bhanu Pratap Singh Rawat, and Hong Yu. 2021 · 2021
Earlier work this paper cites.
Online question and answer sessions: How students support their own and other students’ processes of inquiry in a text-based learning environment
Malin Jansson, Stefan Hrastinski, Stefan Stenbom, and Fredrik Enoksson. 2021 · 2021
Earlier work this paper cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021 · 2021
Earlier work this paper cites.
Backdoor attacks on pre-trained models by layerwise weight poisoning
Linyang Li, Demin Song, Xiaonan Li, Jiehang Zeng, Ruotian Ma, and Xipeng Qiu. 2021b · 2021
Earlier work this paper cites.
Large language models can be strong differentially private learners
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2021c · 2021
Earlier work this paper cites.
Neural attention distillation: Erasing backdoor triggers from deep neural networks
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021a · 2021
Earlier work this paper cites.
What do you see? Evaluation of explainable artificial intelligence (XAI) interpretability through neural backdoors. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining . 1027–1035
Yi-Shan Lin, Wen-Chuan Lee, and Z Berkay Celik. 2021 · 2021
Earlier work this paper cites.
Hidden Killer: Invisible textual backdoor attacks with syntactic trigger
Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun. 2021 · 2021
Earlier work this paper cites.
You autocomplete me: Poisoning vulnerabilities in neural code completion. In 30th USENIX Security Symposium (USENIX Security 21) . 1559–1575
Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. 2021 · 2021
Earlier work this paper cites.
BDDR: An effective defense against textual backdoor attacks
Kun Shao, Junan Yang, Yang Ai, Hui Liu, and Yu Zhang. 2021 · 2021
Earlier work this paper cites.
Membership inference attacks against NLP classification models. In NeurIPS 2021 Workshop Privacy in Machine Learning
Virat Shejwalkar, Huseyin A Inan, Amir Houmansadr, and Robert Sim. 2021 · 2021
Earlier work this paper cites.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Earlier work this paper cites.
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He. 2021a · 2021
Earlier work this paper cites.
RAP: Robustness-aware perturbations for defending against backdoor attacks on NLP models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun. 2021b · 2021
Earlier work this paper cites.
Defense against adversarial attacks by reconstructing images
Shudong Zhang, Haichang Gao, and Qingxun Rao. 2021a · 2021
Earlier work this paper cites.
Trojaning language models for fun and profit. In 2021 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, 179–197
Xinyang Zhang, Zheng Zhang, Shouling Ji, and Ting Wang. 2021b · 2021
Earlier work this paper cites.
Exploiting explanations for model inversion attacks. In Proceedings of the IEEE/CVF international conference on computer vision . 682–692
Xuejun Zhao, Wencan Zhang, Xiaokui Xiao, and Brian Lim. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
LAMP: Extracting text from gradients with language model priors
Mislav Balunovic, Dimitar Dimitrov, Nikola Jovanović, and Martin Vechev. 2022 · 2022
Earlier work this paper cites.
Evaluating the susceptibility of pre-trained language models via handcrafted adversarial examples
Hezekiah J Branch, Jonathan Rodriguez Cefalu, Jeremy McHugh, Leyla Hujer, Aditya Bahl, Daniel del Castillo Iglesias, Ron Heichman, and Ramesh Darwishi. 2022 · 2022
Earlier work this paper cites.
BadPrompt: Backdoor attacks on continuous prompts
Xiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang, et al · 2022
Earlier work this paper cites.
A unified evaluation of textual backdoor learning: Frameworks and benchmarks
Ganqu Cui, Lifan Yuan, Bingxiang He, Yangyi Chen, Zhiyuan Liu, and Maosong Sun. 2022 · 2022
Earlier work this paper cites.
PPT: Backdoor attacks on pre-trained models via poisoned prompt tuning. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 . 680–686
Wei Du, Yichun Zhao, Boqun Li, Gongshen Liu, and Shilin Wang. 2022 · 2022
Earlier work this paper cites.
Demystifying prompts in language models via perplexity estimation
Hila Gonen, Srini Iyer, Terra Blevins, Noah A Smith, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
"Why Meta’s latest large language model survived only three days online"
Will Douglas Heaven. 2022 · 2022
Earlier work this paper cites.
Are Large Pre-Trained Language Models Leaking Your Personal Information?
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. 2022 · 2022
Earlier work this paper cites.
Deduplicating training data mitigates privacy risks in language models. In International Conference on Machine Learning . PMLR, 10697–10707
Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
Behind the Mask: Demographic bias in name detection for PII masking
Courtney Mansfield, Amandalynne Paullada, and Kristen Howell. 2022 · 2022
Earlier work this paper cites.
Quantifying privacy risks of masked language models using membership inference attacks
Fatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, and Reza Shokri. 2022 · 2022
Earlier work this paper cites.
Diffusion models for adversarial purification
Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, and Anima Anandkumar. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Cited alongside, same era.
Backdoors in neural models of source code. In 2022 26th International Conference on Pattern Recognition (ICPR) . IEEE, 2892–2899
Goutham Ramakrishnan and Aws Albarghouthi. 2022 · 2022
Cited alongside, same era.
Fine-tuning is all you need to mitigate backdoor attacks
Zeyang Sha, Xinlei He, Pascal Berrang, Mathias Humbert, and Yang Zhang. 2022 · 2022
Cited alongside, same era.
An information-theoretic approach to prompt engineering without ground truth labels
VistaGPT: Generative parallel transformers for vehicles with intelligent systems for transport automation
Yonglin Tian, Xuan Li, Hui Zhang, Chen Zhao, Bai Li, Xiao Wang, and Fei-Yue Wang. 2023 · 2023
Later among the works it cites.
Privinfer: Privacy-preserving inference for black-box large language model
Meng Tong, Kejiang Chen, Yuang Qi, Jie Zhang, Weiming Zhang, and Nenghai Yu. 2023 · 2023
Later among the works it cites.
AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks
Alexander Tornede, Difan Deng, Theresa Eimer, Joseph Giovanelli, Aditya Mohan, Tim Ruhkopf, Sarah Segel, Daphne Theodorakopoulos, Tanja Tornede, Henning Wachsmuth, et al · 2023
Later among the works it cites.
Poisoning Language Models During Instruction Tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw, Kyle Jeffrey Rogers, Alexia Pauline Delorey, Mahmoud Khalil, Nancy Fulda, and David Wingate. 2022 · 2022
Cited alongside, same era.
Federated vehicular transformers and their federations: Privacy-preserving computing and cooperation for autonomous driving
Yonglin Tian, Jiangong Wang, Yutong Wang, Chen Zhao, Fei Yao, and Xiao Wang. 2022 · 2022
Cited alongside, same era.
Downstream task performance of BERT models pre-trained using automatically de-identified clinical data. In Proceedings of the Thirteenth Language Resources and Evaluation Conference . 4245–4252
Thomas Vakili, Anastasios Lamproudis, Aron Henriksson, and Hercules Dalianis. 2022 · 2022
Cited alongside, same era.
Perplexity from plm is unreliable for evaluating text quality
Yequan Wang, Jiawen Deng, Aixin Sun, and Xuying Meng. 2022a · 2022
Cited alongside, same era.
Analyzing and Defending against Membership Inference Attacks in Natural Language Processing Classification. In 2022 IEEE International Conference on Big Data (Big Data) . IEEE, 5823–5832
Yijue Wang, Nuo Xu, Shaoyi Huang, Kaleel Mahmood, Dan Guo, Caiwen Ding, Wujie Wen, and Sanguthevar Rajasekaran. 2022b · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . 214–229
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al · 2022
Cited alongside, same era.
Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, et al · 2023
Later among the works it cites.
Adversarial Demonstration Attacks on Large Language Models
Jiongxiao Wang, Zichen Liu, Keun Hee Park, Muhao Chen, and Chaowei Xiao. 2023b · 2023
Later among the works it cites.
Self-deception: Reverse penetrating the semantic firewall of large language models
Zhenhua Wang, Wei Xie, Kai Chen, Baosheng Wang, Zhiwen Gui, and Enze Wang. 2023c · 2023
Later among the works it cites.
Rab: Provable robustness against backdoor attacks. In 2023 IEEE Symposium on Security and Privacy (SP) . IEEE, 1311–1328
Maurice Weber, Xiaojun Xu, Bojan Karlaš, Ce Zhang, and Bo Li. 2023 · 2023
Later among the works it cites.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
A prompt pattern catalog to enhance prompt engineering with ChatGPT
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023 · 2023
Later among the works it cites.
Overcoming catastrophic forgetting in massively multilingual continual learning
Genta Indra Winata, Lingjue Xie, Karthik Radhakrishnan, Shijie Wu, Xisen Jin, Pengxiang Cheng, Mayank Kulkarni, and Daniel Preotiuc-Pietro. 2023 · 2023
Later among the works it cites.
Defending ChatGPT against Jailbreak Attack via Self-Reminder
Fangzhao Wu, Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, and Xing Xie. 2023c · 2023
Later among the works it cites.
Unveiling security, privacy, and ethical concerns of ChatGPT
Xiaodong Wu, Ran Duan, and Jianbing Ni. 2023b · 2023
Later among the works it cites.
Selecting and composing learning rate policies for deep neural networks
Yanzhao Wu and Ling Liu. 2023 · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2023
Later among the works it cites.
Defending pre-trained language models as few-shot learners against backdoor attacks
Zhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang, Shouling Ji, Jinghui Chen, Fenglong Ma, and Ting Wang. 2023b · 2023
Later among the works it cites.
Explanation leaks: Explanation-guided model extraction attacks
Anli Yan, Teng Huang, Lishan Ke, Xiaozhang Liu, Qi Chen, and Changyu Dong. 2023 · 2023
Later among the works it cites.
A Comprehensive Overview of Backdoor Attacks in Large Language Models within Communication Networks
Haomiao Yang, Kunlan Xiang, Hongwei Li, and Rongxing Lu. 2023 · 2023
Later among the works it cites.
Wencong You, Zayd Hammoudeh, and Daniel Lowd. 2023 · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing. 2023 · 2023
Later among the works it cites.
Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success
Yiming Zhang and Daphne Ippolito. 2023 · 2023
Later among the works it cites.
Defending large language models against jailbreaking attacks through goal prioritization
Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023 · 2023
Later among the works it cites.
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
Shuai Zhao, Jinming Wen, Luu Anh Tuan, Junbo Zhao, and Jie Fu. 2023a · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Chat-GPT is on the Horizon: Could a Large Language Model be Suitable for Intelligent Traffic Safety Research and Applications?
Abdel-Aty Zheng, Wang Wang, and Ding. 2023c · 2023
Later among the works it cites.
Input Reconstruction Attack against Vertical Federated Large Language Models
Fei Zheng. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
Ou Zheng, Mohamed Abdel-Aty, Dongdong Wang, Chenzhu Wang, and Shengxuan Ding. 2023a · 2023
Later among the works it cites.
Efficiently Measuring the Cognitive Ability of LLMs: An Adaptive Testing Perspective
Yan Zhuang, Qi Liu, Yuting Ning, Weizhe Huang, Rui Lv, Zhenya Huang, Guanhao Zhao, Zheng Zhang, Qingyang Mao, Shijin Wang, et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
Unraveling the Mysteries of Anomaly Detection: The Power of Explainable AI
30dascoding. 2024 · 2024
Closest in time.
LLM Guard - The Security Toolkit for LLM Interactions
Protect AI. 2024 · 2024
Closest in time.
Jailbreak Chat
Alex Albert. 2023 · 2024
Closest in time.
Next-generation intrusion detection systems with LLMs: real-time anomaly detection, explainable AI, and adaptive data generation
Tarek Ali. 2024 · 2024
Closest in time.
California Consumer Privacy Act (CCPA)
ROB BONTA. 2023 · 2024
Closest in time.
XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models
Erik Cambria, Lorenzo Malandri, Fabio Mercorio, Navid Nobani, and Andrea Seveso. 2024 · 2024
Closest in time.
Information Systems Security (INFOSEC)
Computer Security Resource Center. 2023 · 2024
Closest in time.
SocraSynth: Multi-LLM Reasoning with Conditional Statistics
Edward Y Chang. 2023 · 2024
Closest in time.
Humans or llms as the judge? a study on judgement biases
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024a · 2024
Closest in time.
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song, and Bo Li. 2024b · 2024
Closest in time.
Comprehensive assessment of jailbreak attacks against llms
Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang. 2024 · 2024
Closest in time.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Closest in time.
What is data privacy?
CLOUDFLARE. 2023 · 2024
Closest in time.
The Galactica AI model was trained on scientific knowledge – but it spat out alarmingly plausible nonsense
The Conversation. 2022 · 2024
Closest in time.
Bias and unfairness in information retrieval systems: New challenges in the llm era. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 6437–6447
Sunhao Dai, Chen Xu, Shicheng Xu, Liang Pang, Zhenhua Dong, and Jun Xu. 2024 · 2024
Closest in time.
PII and Data Scrubbing
develop.sentry.dev. 2023 · 2024
Closest in time.
Machine learning: What are membership inference attacks?
Ben Dickson. 2021 · 2024
Closest in time.
Privacy Implications of Explainable AI in Data-Driven Systems
Fatima Ezzeddine. 2024 · 2024
Closest in time.
A survey on rag meeting llms: Towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 6491–6501
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024 · 2024
Closest in time.
Exposing privacy gaps: Membership inference attack on preference data for LLM alignment
Qizhang Feng, Siva Rajesh Kasa, Hyokun Yun, Choon Hui Teo, and Sravan Babu Bodapati. 2024b · 2024
Closest in time.
Retrieval-generation synergy augmented large language models. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 11661–11665
Zhangyin Feng, Xiaocheng Feng, Dezhi Zhao, Maojin Yang, and Bing Qin. 2024a · 2024
Closest in time.
Data Poisoning Attacks: A New Attack Vector within AI
Jacob Fox. 2023 · 2024
Closest in time.
What Is Personally Identifiable Information (PII)? Types and Examples
JAKE FRANKENFIELD. 2023 · 2024
Closest in time.
Bias and fairness in large language models: A survey
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024 · 2024
Closest in time.
Gradient-based Adversarial Attacks : An Introduction
Siddhant Haldar. 2023 · 2024
Closest in time.
The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies
Feng He, Tianqing Zhu, Dayong Ye, Bo Liu, Wanlei Zhou, and Philip S Yu. 2024b · 2024
Closest in time.
Data Stealing Attacks against Large Language Models via Backdooring
Jiaming He, Guanyu Hou, Xinyue Jia, Yangyang Chen, Wenqi Liao, Yinhang Zhou, and Rang Zhou. 2024a · 2024
Closest in time.
Mitigating catastrophic forgetting in large language models with self-synthesized rehearsal
Jianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang, Xinting Liao, Linfeng Song, Junfeng Yao, and Jinsong Su. 2024 · 2024
Closest in time.
Attackeval: How to evaluate the effectiveness of jailbreak attacking on large language models
Mingyu Jin, Suiyuan Zhu, Beichen Wang, Zihao Zhou, Chong Zhang, Yongfeng Zhang, et al · 2024
Closest in time.
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
Masahiro Kaneko, Youmi Ma, Yuki Wata, and Naoaki Okazaki. 2024 · 2024
Closest in time.
Propile: Probing privacy leakage in large language models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2024 · 2024
Closest in time.
Isack Lee and Haebin Seong. 2024 · 2024
Closest in time.
Vl-trojan: Multimodal instruction backdoor attacks against autoregressive visual language models
Jiawei Liang, Siyuan Liang, Man Luo, Aishan Liu, Dongchen Han, Ee-Chien Chang, and Xiaochun Cao. 2024 · 2024
Closest in time.
Automatic and universal prompt injection attacks against large language models
Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao. 2024b · 2024
Closest in time.
What is iPhone jailbreaking
Malwarebytes. 2023 · 2024
Closest in time.
Reinventing search with a new AI-powered Microsoft Bing and Edge, your copilot for the web
Yusuf Mehdi. 2023 · 2024
Closest in time.
Llama-2-13b-hf
Meta-Llma. 2023a · 2024
Closest in time.
Llama-2-7b
Meta-Llma. 2023b · 2024
Closest in time.
LLaMA-30b
Meta-Llma. 2023c · 2024
Closest in time.
Large language models: A survey
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024 · 2024
Closest in time.
Prompt Hacking and Misuse of LLMs
Aayush Mittal. 2023 · 2024
Closest in time.
What Is Rooting? The Risks of Rooting Your Android Device
Domenic Molinaro. 2023 · 2024
Closest in time.
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
Krishna Kanth Nakka, Ahmed Frikha, Ricardo Mendes, Xue Jiang, and Xuebing Zhou. 2024 · 2024
Closest in time.
Why root Android phones?
Jomilė Nakutavičiūtė. 2023 · 2024
Closest in time.
Backdoor attacks and defenses in federated learning: Survey, challenges and future research directions
Thuy Dung Nguyen, Tuan Nguyen, Phi Le Nguyen, Hieu H Pham, Khoa D Doan, and Kok-Seng Wong. 2024 · 2024
Closest in time.
API to Prevent Prompt Injection & Jailbreaks
OpenAI. 2023a · 2024
Closest in time.
How do you protect Machine Learning from attacks?
Prachi (Nayyar) Pathak. 2023 · 2024
Closest in time.
WHAT IS JAILBREAKING, CRACKING, OR ROOTING A MOBILE DEVICE?
Pearlhawaii.com. 2023 · 2024
Closest in time.
Methodologies of Jailbreaking
Miguel Piedrafita. 2022 · 2024
Closest in time.
Promptbase
Promptbase. 2024 · 2024
Closest in time.
Jailbreaking
Learn Prompting. 2023a · 2024
Closest in time.
Your Guide to Generative AI
Learn Prompting. 2023b · 2024
Closest in time.
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
Delong Ran, Jinyuan Liu, Yichen Gong, Jingyi Zheng, Xinlei He, Tianshuo Cong, and Anyu Wang. 2024 · 2024
Closest in time.
LAMA: LAnguage Model Analysis
Facebook Resrach. 2024 · 2024
Closest in time.
Recent Applications of Explainable AI (XAI): A Systematic Literature Review
Mirka Saarela and Vili Podgorelec. 2024 · 2024
Closest in time.
Instruction Defense
Sander Schulhoff. 2024a · 2024
Closest in time.
Instruction Defense
Sander Schulhoff. 2024b · 2024
Closest in time.
Sandwich Defense
Sander Schulhoff. 2024c · 2024
Closest in time.
Prompt Hacking and Misuse of LLMs
SECWRITER. 2023 · 2024
Closest in time.
Exploring Prompt Injection Attacks
Jose Selvi. 2023 · 2024
Closest in time.
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. 2024 · 2024
Closest in time.
TITANIC: Towards Production Federated Learning with Large Language Models
Ningxin Su, Chenghao Hu, Baochun Li, and Bo Li. 2023 · 2024
Closest in time.
TrustLLM: Trustworthiness in Large Language Models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Closest in time.
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
Annalisa Szymanski, Noah Ziems, Heather A Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, and Ronald A Metoyer. 2024 · 2024
Closest in time.
Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations
Apostol Vassilev, Alina Oprea, Alie Fordyce, and Hyrum Andersen. 2024 · 2024
Closest in time.
Unique Security and Privacy Threats of Large Language Model: A Comprehensive Survey
Shang Wang, Tianqing Zhu, Bo Liu, Ding Ming, Xu Guo, Dayong Ye, and Wanlei Zhou. 2024b · 2024
Closest in time.
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
Yifei Wang, Dizhan Xue, Shengjie Zhang, and Shengsheng Qian. 2024a · 2024
Closest in time.
Can Large Language Models Automatically Jailbreak GPT-4V?
Yuanwei Wu, Yue Huang, Yixin Liu, Xiang Li, Pan Zhou, and Lichao Sun. 2024 · 2024
Closest in time.
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li. 2024a · 2024
Closest in time.
CBD: A certified backdoor detector based on local dominant probability
Zhen Xiang, Zidi Xiong, and Bo Li. 2024b · 2024
Closest in time.
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
Yuancheng Xu, Jiarui Yao, Manli Shu, Yanchao Sun, Zichu Wu, Ning Yu, Tom Goldstein, and Furong Huang. 2024 · 2024
Closest in time.
ParaFuzz: An interpretability-driven technique for detecting poisoned samples in nlp
Lu Yan, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang, Xuan Chen, Guangyu Shen, and Xiangyu Zhang. 2024b · 2024
Closest in time.
Shenao Yan, Shen Wang, Yue Duan, Hanbin Hong, Kiho Lee, Doowon Kim, and Yuan Hong. 2024a · 2024
Closest in time.
Zhou Yang, Zhensu Sun, Terry Zhuo Yue, Premkumar Devanbu, and David Lo. 2024a · 2024
Closest in time.
Stealthy backdoor attack for code models
Zhou Yang, Bowen Xu, Jie M Zhang, Hong Jin Kang, Jieke Shi, Junda He, and David Lo. 2024b · 2024
Closest in time.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024 · 2024
Closest in time.
Don’t Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron, Chaowei Xiao, and Ning Zhang. 2024 · 2024
Closest in time.
Siyun Zhao, Yuqing Yang, Zilong Wang, Zhiyuan He, Luna K Qiu, and Lili Qiu. 2024 · 2024
Closest in time.
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia, Yingshuang Gu, Mingxu Chai, Fukang Zhu, Caishuang Huang, Shihan Dou, Zhiheng Xi, et al · 2024
Closest in time.
Demystifying membership inference attacks in machine learning as a service
Stacey Truex, Ling Liu, Mehmet Emre Gursoy, Lei Yu, and Wenqi Wei. 2019 · 2089
Closest in time.