Fetching the paper…
Reading the bibliography…
Recent latent-space monitoring techniques have shown promise as defenses against LLM attacks.
Are Odds Really Odd? Bypassing Statistical Detection of Adversarial Examples
Hossein Hosseini, Sreeram Kannan, and Radha Poovendran · 1907
Earlier work this paper cites.
On the generalized distance in Statistics
P. C. Mahalanobis · 1936
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2013
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Adversarial Examples Detection in Deep Networks with Convolutional Filter Statistics
Xin Li and Fuxin Li · 2016
Earlier work this paper cites.
Probing Classifiers: Promises, Shortcomings, and Advances
Yonatan Belinkov · 2017
Earlier work this paper cites.
Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Detecting Adversarial Samples from Artifacts
Reuben Feinman, Ryan R. Curtin, Saurabh Shintre, and Andrew B. Gardner · 2017
Earlier work this paper cites.
On the (Statistical) Detection of Adversarial Examples
Kathrin Grosse, Praveen Manoharan, Nicolas Papernot, Michael Backes, and Patrick McDaniel · 2017
Earlier work this paper cites.
On Detecting Adversarial Perturbations
Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff · 2017
Earlier work this paper cites.
Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
Victor Zhong, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Trace and Detect Adversarial Attacks on CNNs Using Feature Response Maps
Mohammadreza Amirian, Friedhelm Schwenker, and Thilo Stadelmann · 2018
Earlier work this paper cites.
Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
Anish Athalye, Nicholas Carlini, and David Wagner · 2018
Earlier work this paper cites.
sql-create-context Dataset, 2023
b-mc2 · 2018
Earlier work this paper cites.
Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Ben Edwards, Taesung Lee, Ian Molloy, and B. Srivastava · 2018
Earlier work this paper cites.
Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality
Xingjun Ma, Bo Li, Yisen Wang, Sarah Monazam Erfani, Sudanthi N. R. Wijewickrema, Michael E. Houle, Grant Robert Schoenebeck, Dawn Xiaodong Song, and James Bailey · 2018
Earlier work this paper cites.
Spectral Signatures in Backdoor Attacks
Brandon Tran, Jerry Li, and Aleksander Madry · 2018
Earlier work this paper cites.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev · 2018
Earlier work this paper cites.
Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
STRIP: a defence against trojan attacks on deep neural networks
Yansong Gao, Chang Xu, Derui Wang, Shiping Chen, Damith Chinthana Ranasinghe, and Surya Nepal · 2019
Earlier work this paper cites.
Gradient hacking
Evan Hubinger · 2019
Earlier work this paper cites.
Bypassing Backdoor Detection Algorithms in Deep Learning
Te Juin Lester Tan and Reza Shokri · 2020
Earlier work this paper cites.
Adversarial Example Detection Using Latent Neighborhood Graph
Ahmed Abusnaina, Yuhang Wu, Sunpreet Arora, Yizhen Wang, Fei Wang, Hao Yang, and David Mohaisen · 2021
Earlier work this paper cites.
Backdoor Attack with Imperceptible Input and Latent Modification
Khoa Doan, Yingjie Lao, and Ping Li · 2021
Earlier work this paper cites.
SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics
Jonathan Hayase, Weihao Kong, Raghav Somani, and Sewoong Oh · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised Learning
Jinyuan Jia, Yupei Liu, and Neil Zhenqiang Gong · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection
Alexander Podolskiy, Dmitry Lipin, Andrey Bout, Ekaterina Artemova, and Irina Piontkovskaya · 2021
Earlier work this paper cites.
A General Framework For Detecting Anomalous Inputs to DNN Classifiers
Jayaram Raghuram, Varun Chandrasekaran, Somesh Jha, and Suman Banerjee · 2021
Earlier work this paper cites.
Demon in the Variant: Statistical Analysis of DNNs for Robust Backdoor Contamination Detection
Di Tang, XiaoFeng Wang, Haixu Tang, and Kehuan Zhang · 2021
Earlier work this paper cites.
Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He · 2021
Earlier work this paper cites.
Zeyu Yun, Yubei Chen, Bruno A. Olshausen, and Yann LeCun · 2021
Earlier work this paper cites.
Discovering Latent Knowledge in Language Models Without Supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks
Sishuo Chen, Wenkai Yang, Zhiyuan Zhang, Xiaohan Bi, and Xu Sun · 2022
Earlier work this paper cites.
Planting Undetectable Backdoors in Machine Learning Models
Shafi Goldwasser, Michael P. Kim, Vinod Vaikuntanathan, and Or Zamir · 2022
Earlier work this paper cites.
Generating Distributional Adversarial Examples to Evade Statistical Detectors
Yigitcan Kaya, Muhammad Bilal Zafar, Sergul Aydore, Nathalie Rauschmayr, and Krishnaram Kenthapadi · 2022
Cited alongside, same era.
Piccolo: Exposing Complex Backdoors in NLP Transformer Models
Yingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An, Shiqing Ma, and Xiangyu Zhang · 2022
Cited alongside, same era.
Circumventing Backdoor Defenses that are Based on Latent Separability
Xiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar, and Prateek Mittal · 2022
Cited alongside, same era.
Circumventing interpretability: How to defeat mind-readers
Lee Sharkey · 2022
Cited alongside, same era.
A Survey on Backdoor Attack and Defense in Natural Language Processing
Xuan Sheng, Zhaoyang Han, Piji Li, and Xiangmao Chang · 2022
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
David Chanin, James Wilken-Smith, Tomáš Dulka, Hardik Bhatnagar, and Joseph Bloom · 2024
Closest in time.
Model manipulation attacks enable more rigorous evaluations of llm capabilities
Zora Che, Stephen Casper, Anirudh Satheesh, Rohit Gandikota, Domenic Rosati, Stewart Slocum, Lev E McKinney, Zichu Wu, Zikui Cai, Bilal Chughtai, et al · 2024
Closest in time.
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
Joshua Clymer, Caden Juang, and Severin Field · 2024
Closest in time.
Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness
Stanislav Fort and Balaji Lakshminarayanan · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A Subspace Projective Clustering Approach for Backdoor Attack Detection and Mitigation in Deep Neural Networks
Yue Wang, Wenqing Li, Esha Sarkar, Muhammad Shafique, Michail Maniatakos, and Saif Eddin G. Jabari · 2022
Cited alongside, same era.
Eliciting Latent Predictions from Transformers with the Tuned Lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Jailbreaking Black Box Large Language Models in Twenty Queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Cited alongside, same era.
Code Alpaca: An Instruction-following LLaMA model for code generation
Sahil Chaudhary · 2023
Cited alongside, same era.
Sparse Autoencoders Find Highly Interpretable Features in Language Models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou · 2023
Cited alongside, same era.
Rohit Gandikota, Sheridan Feucht, Samuel Marks, and David Bau · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Closest in time.
Coercing LLMs to Do and Reveal (Almost) Anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein · 2024
Closest in time.
Automated Multi-Turn Red-Teaming with Cascade, October 2024
Haize · 2024
Closest in time.
Backdoors as an analogy for deceptive alignment, 2024
Jacob Hilton and Mark Xu · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
Samyak Jain, Ekdeep Singh Lubana, Kemal Oksuz, Tom Joy, Philip H. S. Torr, Amartya Sanyal, and Puneet K. Dokania · 2024
Closest in time.
Haibo Jin, Leyang Hu, Xinuo Li, Peiyan Zhang, Chonghan Chen, Jun Zhuang, and Haohan Wang · 2024
Closest in time.
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
Nathalie Maria Kirch, Severin Field, and Stephen Casper · 2024
Closest in time.
SAEs (usually) Transfer Between Base and Chat Models
Connor Kissane, Robert Krzyzanowski, Arthur Conmy, and Neel Nanda · 2024
Closest in time.
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue · 2024
Closest in time.
BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning
Siyuan Liang, Mingli Zhu, Aishan Liu, Baoyuan Wu, Xiaochun Cao, and Ee-Chien Chang · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Simple probes can catch sleeper agents
Monte MacDiarmid, Timothy Maxwell, Nicholas Schiefer, Jesse Mu, Jared Kaplan, David Duvenaud, Sam Bowman, Alex Tamkin, Ethan Perez, Mrinank Sharma, Carson Denison, and Evan Hubinger · 2024
Closest in time.
Deep Causal Transcoding: A Framework for Mechanistically Eliciting Latent Behaviors in Language Models
Andrew Mack and Alex Turner · 2024
Closest in time.
Samuel Marks and Max Tegmark · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Robust Backdoor Detection for Deep Learning via Topological Evolution Dynamics
Xiaoxing Mo, Yechao Zhang, Leo Yu Zhang, Wei Luo, Nan Sun, Shengshan Hu, Shang Gao, and Yang Xiang · 2024
Closest in time.
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
Sara Price, Arjun Panickssery, Sam Bowman, and Asa Cooper Stickland · 2024
Closest in time.
Representation noising effectively prevents harmful fine-tuning on LLMs
Domenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze, David Atanasov, Robie Gonzales, Subhabrata Majumdar, Carsten Maple, Hassan Sajjad, and Frank Rudzicz · 2024
Closest in time.
Public comment: Robustness evaluation seems invalid, 2024
Christian Schlarmann, Francesco Croce, and Matthias Hein · 2024
Closest in time.
Revisiting the Robust Alignment of Circuit Breakers
Leo Schwinn and Simon Geisler · 2024
Closest in time.
Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, et al · 2024
Closest in time.
A StrongREJECT for empty jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al · 2024
Closest in time.
Analyzing the Generalization and Reliability of Steering Vectors
Daniel Tan, David Chanin, Aengus Lynch, Dimitrios Kanoulas, Brooks Paige, Adria Garriga-Alonso, and Robert Kirk · 2024
Closest in time.
Distribution Preserving Backdoor Attack in Self-supervised Learning
Guanhong Tao, Zhenting Wang, Shiwei Feng, Guangyu Shen, Shiqing Ma, and Xiangyu Zhang · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
FLRT: Fluent Student-Teacher Redteaming
T Ben Thompson and Michael Sklar · 2024
Closest in time.
Efficient Adversarial Training in LLMs with Continuous Attacks
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel, and Leo Schwinn · 2024
Closest in time.
Jailbreak attacks and defenses against large language models: A survey
Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li · 2024
Closest in time.
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
Shaolei Zhang, Tian Yu, and Yang Feng · 2024
Closest in time.
Improving Alignment and Robustness with Circuit Breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Closest in time.
An Adversarial Perspective on Machine Unlearning for AI Safety
Jakub Łucki, Boyi Wei, Yangsibo Huang, Peter Henderson, Florian Tramèr, and Javier Rando · 2024
Closest in time.