Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are vulnerable to adversarial attacks that add malicious tokens to an input prompt to bypass the safety guardrails of an LLM and cause it to produce harmful content.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan · 2010
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Adversarial examples are not easily detected: Bypassing ten detection methods
Nicholas Carlini and David A. Wagner · 2017
Earlier work this paper cites.
Adversarial examples detection in deep networks with convolutional filter statistics
Xin Li and Fuxin Li · 2017
Earlier work this paper cites.
On the (statistical) detection of adversarial examples
Kathrin Grosse, Praveen Manoharan, Nicolas Papernot, Michael Backes, and Patrick D. McDaniel · 2017
Earlier work this paper cites.
Adversarial and clean data are not twins
Zhitao Gong, Wenlu Wang, and Wei-Shinn Ku · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Thermometer encoding: One hot way to resist adversarial examples
Jacob Buckman, Aurko Roy, Colin Raffel, and Ian J. Goodfellow · 2018
Earlier work this paper cites.
Countering adversarial images using input transformations
Chuan Guo, Mayank Rana, Moustapha Cissé, and Laurens van der Maaten · 2018
Earlier work this paper cites.
Stochastic activation pruning for robust adversarial defense
Guneet S. Dhillon, Kamyar Azizzadenesheli, Zachary C. Lipton, Jeremy Bernstein, Jean Kossaifi, Aran Khanna, and Animashree Anandkumar · 2018
Earlier work this paper cites.
Adversarial risk and the dangers of evaluating against weak attacks
Jonathan Uesato, Brendan O’Donoghue, Pushmeet Kohli, and Aäron van den Oord · 2018
Earlier work this paper cites.
On the effectiveness of interval bound propagation for training verifiably robust models, 2018
Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy Mann, and Pushmeet Kohli · 2018
Earlier work this paper cites.
Training verified learners with learned verifiers, 2018
Krishnamurthy Dvijotham, Sven Gowal, Robert Stanforth, Relja Arandjelovic, Brendan O’Donoghue, Jonathan Uesato, and Pushmeet Kohli · 2018
Earlier work this paper cites.
Differentiable abstract interpretation for provably robust neural networks
Matthew Mirman, Timon Gehr, and Martin Vechev · 2018
Earlier work this paper cites.
Provable defenses against adversarial examples via the convex outer adversarial polytope
Eric Wong and J. Zico Kolter · 2018
Earlier work this paper cites.
Semidefinite relaxations for certifying robustness to adversarial examples
Aditi Raghunathan, Jacob Steinhardt, and Percy Liang · 2018
Cited alongside, same era.
Achieving verified robustness to symbol substitutions via interval bound propagation
Po-Sen Huang, Robert Stanforth, Johannes Welbl, Chris Dyer, Dani Yogatama, Sven Gowal, Krishnamurthy Dvijotham, and Pushmeet Kohli · 2019
Cited alongside, same era.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Cited alongside, same era.
Certified robustness to adversarial examples with differential privacy
Mathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu, and Suman Jana · 2019
Cited alongside, same era.
Certified adversarial robustness with additive noise
Bai Li, Changyou Chen, Wenlin Wang, and Lawrence Carin · 2019
Cited alongside, same era.
Provably robust deep learning via adversarially trained smoothed classifiers
Why should adversarial perturbations be imperceptible? rethink the research paradigm in adversarial NLP
Yangyi Chen, Hongcheng Gao, Ganqu Cui, Fanchao Qi, Longtao Huang, Zhiyuan Liu, and Maosong Sun · 2022
Later among the works it cites.
Textual manifold-based defense against natural language adversarial examples
Dang Nguyen Minh and Anh Tuan Luu · 2022
Later among the works it cites.
Detection of adversarial examples in text classification: Benchmark and baseline via robust density estimation
KiYoon Yoo, Jangho Kim, Jiho Jang, and Nojun Kwak · 2022
Later among the works it cites.
Detecting word-level adversarial text attacks via SHapley additive exPlanations
Lukas Huber, Marc Alexander Kühn, Edoardo Mosca, and Georg Groh · 2022
Later among the works it cites.
Certified robustness against natural language attacks by causal intervention
Haiteng Zhao, Chang Ma, Xinshuai Dong, Anh Tuan Luu, Zhi-Hong Deng, and Hanwang Zhang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hadi Salman, Jerry Li, Ilya P. Razenshteyn, Pengchuan Zhang, Huan Zhang, Sébastien Bubeck, and Greg Yang · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
On adaptive attacks to adversarial example defenses
Florian Tramèr, Nicholas Carlini, Wieland Brendel, and Aleksander Madry · 2020
Cited alongside, same era.
Second-order provable defenses against adversarial attacks
Sahil Singla and Soheil Feizi · 2020
Cited alongside, same era.
SAFER: A structure-free approach for certified robustness to adversarial word substitutions
Mao Ye, Chengyue Gong, and Qiang Liu · 2020
Cited alongside, same era.
LAFEAT: piercing through adversarial defenses with latent features
Yunrui Yu, Xitong Gao, and Cheng-Zhong Xu · 2021
Cited alongside, same era.
Policy smoothing for provably robust reinforcement learning
Aounon Kumar, Alexander Levine, and Soheil Feizi · 2022
Later among the works it cites.
CROP: Certifying robust policies for reinforcement learning through functional smoothing
Fan Wu, Linyi Li, Zijian Huang, Yevgeniy Vorobeychik, Ding Zhao, and Bo Li · 2022
Later among the works it cites.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, and Ethan Perez · 2023
Closest in time.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Closest in time.
Baseline defenses for adversarial attacks against aligned language models, 2023
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Closest in time.
Detecting language model attacks with perplexity, 2023
Gabriel Alon and Michael Kamfonas · 2023
Closest in time.
Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2023
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Closest in time.
Autodan: Automatic and interpretable adversarial attacks on large language models, 2023
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Certified robustness for large language models with self-denoising
Zhen Zhang, Guanhua Zhang, Bairu Hou, Wenqi Fan, Qing Li, Sijia Liu, Yang Zhang, and Shiyu Chang · 2023
Closest in time.
RS-del: Edit distance robustness certificates for sequence classifiers via randomized deletion
Zhuoqun Huang, Neil G Marchant, Keane Lucas, Lujo Bauer, Olga Ohrimenko, and Benjamin I. P. Rubinstein · 2023
Closest in time.
Provable robustness for streaming models with a sliding window, 2023
Aounon Kumar, Vinu Sankar Sadasivan, and Soheil Feizi · 2023
Closest in time.