Fetching the paper…
Reading the bibliography…
Like a criminal under investigation, Large Language Models (LLMs) might pretend to be aligned while evaluated and misbehave when they have a good opportunity.
Universal Litmus Patterns: Revealing Backdoor Attacks in CNNs, May 2020
Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash, and Heiko Hoffmann · 1906
Earlier work this paper cites.
Februus: Input Purification Defense Against Trojan Attacks on Deep Neural Network Systems
Bao Gia Doan, Ehsan Abbasnejad, and Damith C. Ranasinghe · 1908
Earlier work this paper cites.
Wenbo Guo, Lun Wang, Xinyu Xing, Min Du, and Dawn Song · 1908
Earlier work this paper cites.
Attention is not not Explanation, September 2019
Sarah Wiegreffe and Yuval Pinter · 1908
Earlier work this paper cites.
Detection of Backdoors in Trained Classifiers Without Access to the Training Set, August 2020
Zhen Xiang, David J. Miller, and George Kesidis · 1908
Earlier work this paper cites.
Detecting AI Trojans Using Meta Neural Analysis, October 2020
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A. Gunter, and Bo Li · 1910
Earlier work this paper cites.
"Truth" Drugs in Interrogation - CSI, 1993
George Bimmerle · 1993
Earlier work this paper cites.
Weight Poisoning Attacks on Pre-trained Models, April 2020
Keita Kurita, Paul Michel, and Graham Neubig · 2004
Earlier work this paper cites.
Cassandra: Detecting Trojaned Networks from Adversarial Perturbations, July 2020
Xiaoyu Zhang, Ajmal Mian, Rohit Gupta, Nazanin Rahnavard, and Mubarak Shah · 2007
Earlier work this paper cites.
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, August 2017
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2017
Earlier work this paper cites.
Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, May 2018
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2018
Earlier work this paper cites.
Resilience of Pruned Neural Network Against Poisoning Attack
Bingyin Zhao and Yingjie Lao · 2018
Earlier work this paper cites.
A Benchmark for Interpretability Methods in Deep Neural Networks, November 2019
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Earlier work this paper cites.
ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation
Yingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma, Yousra Aafer, and Xiangyu Zhang · 2019
Earlier work this paper cites.
Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao · 2019
Earlier work this paper cites.
CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model Behavior, October 2022
Eldar David Abraham, Karel D’Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu · 2022
Earlier work this paper cites.
Scaling Laws for Reward Model Overoptimization, October 2022
Leo Gao, John Schulman, and Jacob Hilton · 2022
Cited alongside, same era.
Piccolo: Exposing Complex Backdoors in NLP Transformer Models
Yingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An, Shiqing Ma, and Xiangyu Zhang · 2022
Cited alongside, same era.
Poisoning Attacks and Defenses on Artificial Intelligence: A Survey, February 2022
Miguel A. Ramirez, Song-Kyoo Kim, Hussam Al Hamadi, Ernesto Damiani, Young-Ji Byon, Tae-Yeon Kim, Chung-Suk Cho, and Chan Yeob Yeun · 2022
Cited alongside, same era.
A Survey of Neural Trojan Attacks and Defenses in Deep Learning, February 2022
Jie Wang, Ghulam Mubashar Hassan, and Naveed Akhtar · 2022
Cited alongside, same era.
Core Views on AI Safety: When, Why, What, and How, 2023
Anthropic · 2023
FIND: A Function Description Benchmark for Evaluating Interpretability Methods, December 2023
Sarah Schwettmann, Tamar Rott Shaham, Joanna Materzynska, Neil Chowdhury, Shuang Li, Jacob Andreas, David Bau, and Antonio Torralba · 2023
Later among the works it cites.
Model evaluation for extreme risks, September 2023
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano, and Allan Dafoe · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models, February 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Self-Instruct: Aligning Language Models with Self-Generated Instructions, May 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The Internal State of an LLM Knows When It’s Lying, October 2023
Amos Azaria and Tom Mitchell · 2023
Cited alongside, same era.
Eliciting Latent Predictions from Transformers with the Tuned Lens, November 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
Scheming AIs: Will AIs fake alignment during training in order to get power?, November 2023
Joe Carlsmith · 2023
Cited alongside, same era.
FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions, October 2023
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, and Maarten Sap · 2023
Cited alongside, same era.
Some high-level thoughts on the DeepMind alignment team’s strategy, 2023
Victoria Krakovna and Rohin Sha · 2023
Cited alongside, same era.
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, October 2023
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Cited alongside, same era.
Samuel Marks and Max Tegmark · 2023
Cited alongside, same era.
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen · 2023
Later among the works it cites.
Representation Engineering: A Top-Down Approach to AI Transparency, October 2023
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Later among the works it cites.
OpenXAI: Towards a Transparent Evaluation of Model Explanations, March 2024
Chirag Agarwal, Dan Ley, Satyapriya Krishna, Eshika Saxena, Martin Pawelczyk, Nari Johnson, Isha Puri, Marinka Zitnik, and Himabindu Lakkaraju · 2024
Closest in time.
Simple probes can catch sleeper agents, 2024
Anthropic · 2024
Closest in time.
Discovering Latent Knowledge in Language Models Without Supervision, March 2024
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2024
Closest in time.
Benchmarking Interpretability, 2024
Stephen Casper · 2024
Closest in time.
Black-Box Access is Insufficient for Rigorous AI Audits, January 2024
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger, and Dylan Hadfield-Menell · 2024
Closest in time.
AI Control: Improving Safety Despite Intentional Subversion, January 2024
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger · 2024
Closest in time.
Deception Abilities Emerged in Large Language Models, February 2024
Thilo Hagendorff · 2024
Closest in time.
Linearity of Relation Decoding in Transformer Language Models, February 2024
Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training, January 2024
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
Steering Llama 2 via Contrastive Activation Addition, March 2024
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.