Fetching the paper…
Reading the bibliography…
Recent studies have revealed the vulnerability of large language models to adversarial attacks, where adversaries craft specific input sequences to induce harmful, violent, private, or incorrect outputs.
Towards a robust deep neural network in texts: A survey
Wenqi Wang, Run Wang, Lina Wang, Zhibo Wang, and Aoshuang Ye · 1902
Earlier work this paper cites.
The design and analysis of computer algorithms
Alfred V Aho and John E Hopcroft · 1974
Earlier work this paper cites.
Reversibility and stochastic networks
Frank P Kelly · 2011
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich · 2015
Earlier work this paper cites.
Formal guarantees on the robustness of a classifier against adversarial manipulation
Matthias Hein and Maksym Andriushchenko · 2017
Earlier work this paper cites.
Reluplex: An efficient smt solver for verifying deep neural networks
Guy Katz, Clark Barrett, David L Dill, Kyle Julian, and Mykel J Kochenderfer · 2017
Earlier work this paper cites.
Practical black-box attacks against machine learning
Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z Berkay Celik, and Ananthram Swami · 2017
Earlier work this paper cites.
Deepxplore: Automated whitebox testing of deep learning systems
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana · 2017
Earlier work this paper cites.
On the robustness of the cvpr 2018 white-box adversarial example defenses
Anish Athalye and Nicholas Carlini · 2018
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner · 2018
Earlier work this paper cites.
On adversarial examples for character-level neural machine translation
Javid Ebrahimi, Daniel Lowd, and Dejing Dou · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Towards fast computation of certified robustness for relu networks
Lily Weng, Huan Zhang, Hongge Chen, Zhao Song, Cho-Jui Hsieh, Luca Daniel, Duane Boning, and Inderjit Dhillon · 2018
Earlier work this paper cites.
Spatially transformed adversarial examples
Chaowei Xiao, Jun-Yan Zhu, Bo Li, Warren He, Mingyan Liu, and Dawn Song · 2018
Earlier work this paper cites.
On evaluating adversarial robustness
Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
Efficient and accurate estimation of lipschitz constants for deep neural networks
Mahyar Fazlyab, Alexander Robey, Hamed Hassani, Manfred Morari, and George Pappas · 2019
Earlier work this paper cites.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang · 2019
Earlier work this paper cites.
Tight certificates of adversarial robustness for randomly smoothed classifiers
Guang-He Lee, Yang Yuan, Shiyu Chang, and Tommi Jaakkola · 2019
Earlier work this paper cites.
Provably robust deep learning via adversarially trained smoothed classifiers
Hadi Salman, Jerry Li, Ilya Razenshteyn, Pengchuan Zhang, Huan Zhang, Sebastien Bubeck, and Greg Yang · 2019
Earlier work this paper cites.
Word-level textual adversarial attacking as combinatorial optimization
Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, and Maosong Sun · 2019
Earlier work this paper cites.
Theoretically principled trade-off between robustness and accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan · 2019
Earlier work this paper cites.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Francesco Croce and Matthias Hein · 2020
Earlier work this paper cites.
The complexity of adversarially robust proper learning of halfspaces with agnostic noise
Ilias Diakonikolas, Daniel M Kane, and Pasin Manurangsi · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits · 2020
Earlier work this paper cites.
Robustness certificates for sparse adversarial attacks by randomized ablation
Alexander Levine and Soheil Feizi · 2020
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
John X Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi · 2020
Earlier work this paper cites.
Generating natural language adversarial examples on a large scale with generative models
Yankun Ren, Jianbin Lin, Siliang Tang, Jun Zhou, Shuang Yang, Yuan Qi, and Xiang Ren · 2020
Earlier work this paper cites.
\ \backslash e l l _ 1 ell\_1 adversarial robustness certificates: a randomized smoothing approach
Jiaye Teng, Guang-He Lee, and Yang Yuan · 2020
Earlier work this paper cites.
Safer: A structure-free approach for certified robustness to adversarial word substitutions
Mao Ye, Chengyue Gong, and Qiang Liu · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
On the hardness of robust classification
Pascale Gourdeau, Varun Kanade, Marta Kwiatkowska, and James Worrell · 2021
Cited alongside, same era.
Improved, deterministic smoothing for l_1 certified robustness
Alexander J Levine and Soheil Feizi · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Cited alongside, same era.
Certified robustness to word substitution attack with differential privacy
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
Jindong Wang, HU Xixu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Wei Ye, Haojun Huang, Xiubo Geng, et al · 2023
Later among the works it cites.
Defending chatgpt against jailbreak attack via self-reminder
Fangzhao Wu, Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, and Xing Xie · 2023
Later among the works it cites.
Densepure: Understanding diffusion models for adversarial robustness
Chaowei Xiao, Zhongzhu Chen, Kun Jin, Jiongxiao Wang, Weili Nie, Mingyan Liu, Anima Anandkumar, Bo Li, and Dawn Song · 2023
Later among the works it cites.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenjie Wang, Pengfei Tang, Jian Lou, and Li Xiong · 2021
Cited alongside, same era.
A continuous time framework for discrete denoising models
Andrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth, George Deligiannidis, and Arnaud Doucet · 2022
Cited alongside, same era.
Introduction to algorithms
Thomas H Cormen, Charles E Leiserson, Ronald L Rivest, and Clifford Stein · 2022
Cited alongside, same era.
On the limitations of stochastic pre-processing defenses
Yue Gao, Ilia Shumailov, Kassem Fawaz, and Nicolas Papernot · 2022
Cited alongside, same era.
Text adversarial attacks and defenses: Issues, taxonomy, and perspectives
Xu Han, Ying Zhang, Wei Wang, and Bin Wang · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine · 2022
Cited alongside, same era.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2023
Later among the works it cites.
Certified robustness to text adversarial attacks by randomized [mask]
Jiehang Zeng, Jianhan Xu, Xiaoqing Zheng, and Xuanjing Huang · 2023
Later among the works it cites.
{ \{ DiffSmooth } \} : Certifiably robust learning via diffusion models and local smoothing
Jiawei Zhang, Zhongzhu Chen, Huan Zhang, Chaowei Xiao, and Bo Li · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Detecting language model attacks with perplexity, 2024
Gabriel Alon and Michael J Kamfonas · 2024
Later among the works it cites.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Later among the works it cites.
Gasp: Efficient black-box generation of adversarial suffixes for jailbreaking llms
Advik Raj Basani and Xiao Zhang · 2024
Later among the works it cites.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al · 2024
Later among the works it cites.
The lipschitz-variance-margin tradeoff for enhanced randomized smoothing
Blaise Delattre, Alexandre Araujo, Quentin Barthélemy, and Alexandre Allauzen · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Improved large language model jailbreak detection via pretrained embeddings
Erick Galinkin and Martin Sablotny · 2024
Later among the works it cites.
Jailbreaking proprietary large language models using word substitution cipher
Divij Handa, Advait Chirmule, Bimal Gajera, and Chitta Baral · 2024
Later among the works it cites.
Diffusion denoising as a certified defense against clean-label poisoning
Sanghyun Hong, Nicholas Carlini, and Alexey Kurakin · 2024
Later among the works it cites.
Zhanhao Hu, Julien Piet, Geng Zhao, Jiantao Jiao, and David Wagner · 2024
Later among the works it cites.
Jailbreak chat - the latest in ai chatbot tinkering
Jailbreak Chat · 2024
Later among the works it cites.
Improved techniques for optimization-based jailbreaking on large language models
Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang, Jindong Gu, Yang Liu, Xiaochun Cao, and Min Lin · 2024
Later among the works it cites.
Diffattack: Evasion attacks against diffusion-based adversarial purification
Mintong Kang, Dawn Song, and Bo Li · 2024
Later among the works it cites.
Vishal Kumar, Zeyi Liao, Jaylen Jones, and Huan Sun · 2024
Later among the works it cites.
Zeyi Liao and Huan Sun · 2024
Later among the works it cites.
Fight back against jailbreaking via prompt adversarial tuning
Yichuan Mo, Yuji Wang, Zeming Wei, and Yisen Wang · 2024
Later among the works it cites.
Advprompter: Fast adaptive adversarial prompting for llms
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Later among the works it cites.
Revisiting character-level adversarial attacks
Elias Abad Rocamora, Yongtao Wu, Fanghui Liu, Grigorios G Chrysos, and Volkan Cevher · 2024
Later among the works it cites.
" do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2024
Later among the works it cites.
Solving olympiad geometry without human demonstrations
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong · 2024
Later among the works it cites.
A theoretical understanding of self-correction through in-context alignment
Yifei Wang, Yuyang Wu, Zeming Wei, Stefanie Jegelka, and Yisen Wang · 2024
Later among the works it cites.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Later among the works it cites.
Revisiting jailbreaking for large language models: A representation engineering perspective
Tianlong Li, Zhenghua Wang, Wenhao Liu, Muling Wu, Shihan Dou, Changze Lv, Xiaohua Wang, Xiaoqing Zheng, and Xuan-Jing Huang · 2025
Closest in time.
Boosting jailbreak attack with momentum
Yihao Zhang and Zeming Wei · 2025
Closest in time.