Fetching the paper…
Reading the bibliography…
Despite extensive diagnostics and debugging by developers, AI systems sometimes exhibit harmful unintended behaviors.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Unrestricted adversarial examples
Tom B Brown, Nicholas Carlini, Chiyuan Zhang, Catherine Olsson, Paul Christiano, and Ian Goodfellow · 2018
Earlier work this paper cites.
Regularizing deep networks using efficient layerwise adversarial training
Swami Sankaranarayanan, Arpit Jain, Rama Chellappa, and Ser Nam Lim · 2018
Earlier work this paper cites.
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2018
Earlier work this paper cites.
Worst-case guarantees, 2019
Paul Christiano · 2019
Earlier work this paper cites.
Relaxed adversarial training, Sept 2019
Evan Hubinger · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Earlier work this paper cites.
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao · 2019
Earlier work this paper cites.
Testing robustness against unforeseen adversaries
Max Kaufmann, Daniel Kang, Yi Sun, Steven Basart, Xuwang Yin, Mantas Mazeika, Akul Arora, Adam Dziedzic, Franziska Boenisch, Tom Brown, et al · 2019
Earlier work this paper cites.
Harnessing the vulnerability of latent layers in adversarially trained models, 2019
Mayank Singh, Abhishek Sinha, Nupur Kumari, Harshitha Machiraju, Balaji Krishnamurthy, and Vineeth N Balasubramanian · 2019
Earlier work this paper cites.
Theoretically principled trade-off between robustness and accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan · 2019
Earlier work this paper cites.
Freelb: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu · 2019
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Earlier work this paper cites.
Perceptual adversarial robustness: Defense against unseen threat models
Cassidy Laidlaw, Sahil Singla, and Soheil Feizi · 2020
Earlier work this paper cites.
Adversarial training for large neural language models
Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao · 2020
Earlier work this paper cites.
A closer look at accuracy vs. robustness
Yao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Russ R Salakhutdinov, and Kamalika Chaudhuri · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Earlier work this paper cites.
Pengcheng He, Jianfeng Gao, and Weizhu Chen · 2021
Earlier work this paper cites.
Token-aware virtual adversarial training in natural language understanding
Linyang Li and Xipeng Qiu · 2021
Earlier work this paper cites.
Reliably fast adversarial training via latent adversarial perturbation
Geon Yeong Park and Sang Wan Lee · 2021
Earlier work this paper cites.
Towards speeding up adversarial training in latent spaces
Yaguan Qian, Qiqi Shao, Tengteng Yao, Bin Wang, Shouling Ji, Shaoning Zeng, Zhaoquan Gu, and Wassim Swaileh · 2021
Earlier work this paper cites.
Effect of scale on catastrophic forgetting in neural networks
Vinay Venkatesh Ramasesh, Aitor Lewkowycz, and Ethan Dyer · 2021
Earlier work this paper cites.
Natural language adversarial defense through synonym encoding
Xiaosen Wang, Jin Hao, Yichen Yang, and Kun He · 2021
Earlier work this paper cites.
Removing adversarial noise in class activation feature space
Dawei Zhou, Nannan Wang, Chunlei Peng, Xinbo Gao, Xiaoyu Wang, Jun Yu, and Tongliang Liu · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Earlier work this paper cites.
Continual pre-training mitigates forgetting in language and vision
Andrea Cossu, Tinne Tuytelaars, Antonio Carta, Lucia Passaro, Vincenzo Lomonaco, and Davide Bacciu · 2022
Earlier work this paper cites.
Formulating robustness against unforeseen attacks
Sihui Dai, Saeed Mahloujifar, and Prateek Mittal · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
An overview of backdoor attacks against deep neural networks and possible defences
Wei Guo, Benedetta Tondi, and Mauro Barni · 2022
Earlier work this paper cites.
Latent adversarial training, June 2022
Adam Jermyn · 2022
Earlier work this paper cites.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2022
Cited alongside, same era.
Technical report for iccv 2021 challenge sslad-track3b: Transformers are better continual learners
Duo Li, Guimei Cao, Yunlu Xu, Zhanzhan Cheng, and Yi Niu · 2022
Cited alongside, same era.
Textual manifold-based defense against natural language adversarial examples
Dang Nguyen Minh and Luu Anh Tuan · 2022
Cited alongside, same era.
Improved text classification via contrastive adversarial training
Lin Pan, Chung-Wei Hang, Avirup Sil, and Saloni Potdar · 2022
Cited alongside, same era.
Weighted token-level virtual adversarial training in text classification
Teerapong Sae-Lim and Suronapee Phoomvuthisarn · 2022
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fine-tuned language models are continual learners
Thomas Scialom, Tuhin Chakrabarty, and Smaranda Muresan · 2022
Cited alongside, same era.
Backdoorbench: A comprehensive benchmark of backdoor learning
Baoyuan Wu, Hongrui Chen, Mingda Zhang, Zihao Zhu, Shaokui Wei, Danni Yuan, Chao Shen, and Hongyuan Zha · 2022
Cited alongside, same era.
Adversarial training methods for deep learning: A systematic review
Weimin Zhao, Sanaa Alwidian, and Qusay H Mahmoud · 2022
Cited alongside, same era.
Adversarial training for high-stakes reliability, 2022
Daniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, Buck Shlegeris, and Nate Thomas · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Cited alongside, same era.
Image hijacks: Adversarial images can control generative models at runtime
Luke Bailey, Euan Ong, Stuart Russell, and Scott Emmons · 2023
Cited alongside, same era.
Haoran Wang and Kai Shu · 2023
Later among the works it cites.
Understanding deep representation learning via layerwise feature compression and discrimination
Peng Wang, Xiao Li, Can Yaras, Zhihui Zhu, Laura Balzano, Wei Hu, and Qing Qu · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang · 2023
Later among the works it cites.
Adversarial machine learning in latent representations of neural networks
Milin Zhang, Mohammad Abdi, and Francesco Restuccia · 2023
Later among the works it cites.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Closest in time.
Many-shot jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan J Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, Jamie Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomasz Korbak, Jared Kaplan, Deep Ganguli, Samuel R. Bowman, Ethan Perez, Roger Grosse, and David Duvenaud · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al · 2024
Closest in time.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei W Koh, Daphne Ippolito, Florian Tramer, and Ludwig Schmidt · 2024
Closest in time.
Play guessing game with llm: Indirect jailbreak attack with implicit clues
Zhiyuan Chang, Mingyang Li, Yi Liu, Junjie Wang, Qing Wang, and Yang Liu · 2024
Closest in time.
Sophon: Non-fine-tunable learning to restrain task transferability for pre-trained models
Jiangyi Deng, Shengyuan Pang, Yanjiao Chen, Liangming Xia, Yijie Bai, Haiqin Weng, and Wenyuan Xu · 2024
Closest in time.
Artificial Intelligence Act, April 2024
EU · 2024
Closest in time.
Coercing llms to do and reveal (almost) anything, 2024
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein · 2024
Closest in time.
Attacking large language models with projected gradient descent, 2024
Simon Geisler, Tom Wollschläger, M. H. I. Abdalla, Johannes Gasteiger, and Stephan Günnemann · 2024
Closest in time.
Shashwat Goel, Ameya Prabhu, Philip Torr, Ponnurangam Kumaraguru, and Amartya Sanyal · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Closest in time.
Artprompt: Ascii art-based jailbreak attacks against aligned llms
Fengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang, Bhaskar Ramasubramanian, Bo Li, and Radha Poovendran · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K Kummerfeld, and Rada Mihalcea · 2024
Closest in time.
Investigating bias representations in llama 2 chat via activation steering, 2024
Dawn Lu and Nina Rimsky · 2024
Closest in time.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Safety alignment should be made more than just a few tokens deep
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson · 2024
Closest in time.
Representation noising: A defence mechanism against harmful finetuning
Domenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze, Robie Gonzales, Subhabrata Majumdar, Hassan Sajjad, Frank Rudzicz, et al · 2024
Closest in time.
Soft prompt threats: Attacking safety alignment and unlearning in open-source llms through the embedding space, 2024
Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel, and Stephan Gunnemann · 2024
Closest in time.
Tamper-resistant safeguards for open-weight llms
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, et al · 2024
Closest in time.
A language model’s guide through latent space, 2024
Dimitri von Rütte, Sotiris Anagnostidis, Gregor Bachmann, and Thomas Hofmann · 2024
Closest in time.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson · 2024
Closest in time.
Efficient adversarial training in llms with continuous attacks
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel, and Leo Schwinn · 2024
Closest in time.
Spectral regularization for adversarially-robust representation learning
Sheng Yang, Jacob A Zavatone-Veth, and Cengiz Pehlevan · 2024
Closest in time.
International Scientific Report on the Safety of Advanced AI
Bengio Yohsua, Privitera Daniel, Besiroglu Tamay, Bommasani Rishi, Casper Stephen, Choi Yejin, Goldfarb Danielle, Heidari Hoda, Khalatbari Leila, Longpre Shayne, et al · 2024
Closest in time.
A survey of backdoor attacks and defenses on large language models: Implications for security measures
Shuai Zhao, Meihuizi Jia, Zhongliang Guo, Leilei Gan, Xiaoyu Xu, Xiaobao Wu, Jie Fu, Yichao Feng, Fengjun Pan, and Luu Anh Tuan · 2024
Closest in time.
The elicitation game: Evaluating capability elicitation techniques
Felix Hofstätter, Teun van der Weij, Jayden Teoh, Henning Bartsch, and Francis Rhys Ward · 2025
Closest in time.