Fetching the paper…
Reading the bibliography…
Despite efforts to align large language models (LLMs) with human intentions, widely-used LLMs such as GPT, Llama, and Claude are susceptible to jailbreaking attacks, wherein an adversary fools a targeted LLM into generating objectionable content.
Robustness and accuracy tradeoffs for recommender systems under attack
Carlos E Seminario and David C Wilson · 2012
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Šrndić, Pavel Laskov, Giorgio Giacinto, and Fabio Roli · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
The ai alignment problem: why it is hard, and where to start
Eliezer Yudkowsky · 2016
Earlier work this paper cites.
Adversarial training methods for semi-supervised text classification
Takeru Miyato, Andrew M Dai, and Ian Goodfellow · 2016
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Generalizing to unseen domains via adversarial data augmentation
Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C Duchi, Vittorio Murino, and Silvio Savarese · 2018
Earlier work this paper cites.
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2018
Earlier work this paper cites.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Certified robustness to adversarial examples with differential privacy
Mathias Lecuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu, and Suman Jana · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
Provably robust deep learning via adversarially trained smoothed classifiers
Hadi Salman, Jerry Li, Ilya Razenshteyn, Pengchuan Zhang, Huan Zhang, Sebastien Bubeck, and Greg Yang · 2019
Earlier work this paper cites.
Adversarial training for free!
Ali Shafahi, Mahyar Najibi, Mohammad Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein · 2019
Earlier work this paper cites.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Earlier work this paper cites.
Theoretically principled trade-off between robustness and accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan · 2019
Earlier work this paper cites.
ℓ 1 \ell_{1} adversarial robustness certificates: a randomized smoothing approach
Jiaye Teng, Guang-He Lee, and Yang Yuan · 2019
Earlier work this paper cites.
Generating natural language adversarial examples through probability weighted word saliency
Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che · 2019
Earlier work this paper cites.
Natural language adversarial attack and defense in word level
Xiaosen Wang, Hao Jin, and Kun He · 2019
Earlier work this paper cites.
Combating adversarial misspellings with robust word recognition
Danish Pruthi, Bhuwan Dhingra, and Zachary C Lipton · 2019
Earlier work this paper cites.
Xinyu Zhang, Qiang Wang, Jian Zhang, and Zhao Zhong · 2019
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
The alignment problem: Machine learning and human values
Brian Christian · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Adversarial training for large neural language models
Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao · 2020
Earlier work this paper cites.
On adaptive attacks to adversarial example defenses
Florian Tramer, Nicholas Carlini, Wieland Brendel, and Aleksander Madry · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al · 2020
Earlier work this paper cites.
Robustbench: a standardized adversarial robustness benchmark
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein · 2020
Cited alongside, same era.
Fast is better than free: Revisiting adversarial training
Eric Wong, Leslie Rice, and J Zico Kolter · 2020
Cited alongside, same era.
Denoised smoothing: A provable defense for pretrained classifiers
Hadi Salman, Mingjie Sun, Greg Yang, Ashish Kapoor, and J Zico Kolter · 2020
Cited alongside, same era.
Precise tradeoffs in adversarial training for linear regression
Adel Javanmard, Mahdi Soltanolkotabi, and Hamed Hassani · 2020
Cited alongside, same era.
Perceptual adversarial robustness: Defense against unseen threat models
Cassidy Laidlaw, Sahil Singla, and Soheil Feizi · 2020
Evaluating the adversarial robustness of adaptive test-time defenses
Francesco Croce, Sven Gowal, Thomas Brunner, Evan Shelhamer, Matthias Hein, and Taylan Cemgil · 2022
Later among the works it cites.
Regulating chatgpt and other large generative ai models
Philipp Hacker, Andreas Engel, and Marco Mauer · 2023
Closest in time.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model-based robust deep learning: Generalizing to natural, out-of-distribution data
Alexander Robey, Hamed Hassani, and George J Pappas · 2020
Cited alongside, same era.
Learning perturbation sets for robust machine learning
Eric Wong and J Zico Kolter · 2020
Cited alongside, same era.
Breeds: Benchmarks for subpopulation shift
Shibani Santurkar, Dimitris Tsipras, and Aleksander Madry · 2020
Cited alongside, same era.
Randomized smoothing of all shapes and sizes
Greg Yang, Tony Duan, J Edward Hu, Hadi Salman, Ilya Razenshteyn, and Jerry Li · 2020
Cited alongside, same era.
(de) randomized smoothing for certifiable defense against patch attacks
Alexander Levine and Soheil Feizi · 2020
Cited alongside, same era.
Certified defense to image transformations via randomized smoothing
Marc Fischer, Maximilian Baader, and Martin Vechev · 2020
Cited alongside, same era.
Certified robustness to label-flipping attacks via randomized smoothing
Elan Rosenfeld, Ezra Winston, Pradeep Ravikumar, and Zico Kolter · 2020
Cited alongside, same era.
Closest in time.
Adversarial demonstration attacks on large language models
Jiongxiao Wang, Zichen Liu, Keun Hee Park, Muhao Chen, and Chaowei Xiao · 2023
Closest in time.
Chatgpt utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns
Malik Sallam · 2023
Closest in time.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Closest in time.
Adversarial prompting for black box foundation models
Natalie Maus, Patrick Chao, Eric Wong, and Jacob Gardner · 2023
Closest in time.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Closest in time.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Closest in time.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach · 2023
Closest in time.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Closest in time.
Catastrophic jailbreak of open-source llms via exploiting generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen · 2023
Closest in time.
A survey of adversarial defenses and robustness in nlp
Shreya Goyal, Sumanth Doddapaneni, Mitesh M Khapra, and Balaraman Ravindran · 2023
Closest in time.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Closest in time.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju · 2023
Closest in time.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas · 2023
Closest in time.
Instruction-following evaluation for large language models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
Provable tradeoffs in adversarially robust classification
Edgar Dobriban, Hamed Hassani, David Hong, and Alexander Robey · 2023
Closest in time.
Stability guarantees for feature attributions with multiplicative smoothing
Anton Xue, Rajeev Alur, and Eric Wong · 2023
Closest in time.
A safe harbor for ai evaluation and red teaming
Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng-Xin Yong, Suhas Kotha, et al · 2024
Closest in time.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Closest in time.
Zeyi Liao and Huan Sun · 2024
Closest in time.
Attacking large language models with projected gradient descent
Simon Geisler, Tom Wollschläger, MHI Abdalla, Johannes Gasteiger, and Stephan Günnemann · 2024
Closest in time.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al · 2024
Closest in time.