Fetching the paper…
Reading the bibliography…
Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when given the opportunity.
A backdoor attack against lstm-based text classification systems
Jiazhu Dai, Chuanshuai Chen, and Yike Guo · 1905
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems, 2019
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Universal litmus patterns: Revealing backdoor attacks in cnns
Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash, and Heiko Hoffmann · 1906
Earlier work this paper cites.
Models, reasoning and inference
Judea Pearl et al · 2000
Earlier work this paper cites.
The trojai software framework: An opensource tool for embedding trojans into deep learning models
Kiran Karra, Chace Ashcraft, and Neil Fendley · 2003
Earlier work this paper cites.
Blind backdoors in deep learning models
Eugene Bagdasaryan and Vitaly Shmatikov · 2005
Earlier work this paper cites.
Badnl: Backdoor attacks against NLP models
Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, and Yang Zhang · 2006
Earlier work this paper cites.
Trojaning language models for fun and profit
Xinyang Zhang, Zheng Zhang, and Ting Wang · 2008
Earlier work this paper cites.
What you see may not be what you get: Relationships among self-presentation tactics and ratings of interview and job performance
Murray R Barrick, Jonathan A Shaffer, and Sandra W DeGrassi · 2009
Earlier work this paper cites.
Meta-analysis of data from animal studies: A practical guide
H.M. Vesterinen, E.S. Sena, K.J. Egan, T.C. Hirst, L. Churolov, G.L. Currie, A. Antonic, D.W. Howells, and M.R. Macleod · 2013
Earlier work this paper cites.
The Volkswagen scandal
Britt Blackwelder, Katerine Coleman, Sara Colunga-Santoyo, Jeffrey S Harrison, and Danielle Wozniak · 2016
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Backdoor attacks against learning systems
Yujie Ji, Xinyang Zhang, and Ting Wang · 2017
Earlier work this paper cites.
Neural trojans
Yuntao Liu, Yang Xie, and Ankur Srivastava · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2018
Earlier work this paper cites.
Fine-pruning: Defending against backdooring attacks on deep neural networks
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Perez, Chico Q Camargo, and Ard A Louis · 2018
Earlier work this paper cites.
Worst-case guarantees, 2019
Paul Christiano · 2019
Earlier work this paper cites.
ABS: Scanning neural networks for back-doors by artificial brain stimulation
Yingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma, Yousra Aafer, and Xiangyu Zhang · 2019
Earlier work this paper cites.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao · 2019
Earlier work this paper cites.
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, and D. Sculley · 2020
Earlier work this paper cites.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Cited alongside, same era.
LogiQA: A challenge dataset for machine reading comprehension with logical reasoning, 2020
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang · 2020
Cited alongside, same era.
ONION: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2020
Cited alongside, same era.
ConFoc: Content-focus protection against trojan attacks on neural networks
Miguel Villarreal-Vasquez and Bharat Bhargava · 2020
Cited alongside, same era.
Asleep at the keyboard? Assessing the security of GitHub Copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri · 2022
Later among the works it cites.
Learning by distilling context, 2022
Charlie Snell, Dan Klein, and Ruiqi Zhong · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models, 2022
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Taken out of context: On measuring situational awareness in llms, 2023
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans · 2023
Later among the works it cites.
Scheming AIs: Will AIs fake alignment during training in order to get power?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaidi Xu, Sijia Liu, Pin-Yu Chen, Pu Zhao, and Xue Lin · 2020
Cited alongside, same era.
Model Organisms
Rachel A. Ankeny and Sabina Leonelli · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment, 2021
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Cited alongside, same era.
T-Miner: A generative approach to defend against trojan attacks on { \{ DNN-based } \} text classification
Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K Reddy, and Bimal Viswanath · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Neural attention distillation: Erasing backdoor triggers from deep neural networks
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma · 2021
Cited alongside, same era.
Hidden killer: Invisible textual backdoor attacks with syntactic trigger
Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun · 2021
Cited alongside, same era.
You autocomplete me: Poisoning vulnerabilities in neural code completion
Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov · 2021
Cited alongside, same era.
Joe Carlsmith · 2023
Later among the works it cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Later among the works it cites.
Pengzhou Cheng, Zongru Wu, Wei Du, and Gongshen Liu · 2023
Later among the works it cites.
Model organisms of misalignment: The case for a new pillar of alignment research, 2023
Evan Hubinger, Nicholas Schiefer, Carson Denison, and Ethan Perez · 2023
Later among the works it cites.
Towards a situational awareness benchmark for LLMs
Rudolf Laine, Alexander Meinke, and Owain Evans · 2023
Later among the works it cites.
Measuring faithfulness in chain-of-thought reasoning, 2023
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez · 2023
Later among the works it cites.
The trojan detection challenge 2023 (llm edition)
Mantas Mazeika, Andy Zou, Norman Mu, Long Phan, Zifan Wang, Chunru Yu, Adam Khoja, Fengqing Jiang, Aidan O’Gara, Ellie Sakhaee, Zhen Xiang, Arezoo Rajabi, Dan Hendrycks, Radha Poovendran, Bo Li, and David Forsyth · 2023
Later among the works it cites.
GPT-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
AI deception: A survey of examples, risks, and potential solutions
Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback, 2023
Javier Rando and Florian Tramèr · 2023
Later among the works it cites.
Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn · 2023
Later among the works it cites.
Role play with large language models
Murray Shanahan, Kyle McDonell, and Laria Reynolds · 2023
Later among the works it cites.
On the exploitability of instruction tuning, 2023
Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein · 2023
Later among the works it cites.
Uncovering mesa-optimization algorithms in transformers
Johannes von Oswald, Eyvind Niklasson, Maximilian Schlegel, Seijin Kobayashi, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Blaise Agüera y Arcas, Max Vladymyrov, Razvan Pascanu, and João Sacramento · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Later among the works it cites.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen · 2023
Later among the works it cites.
RAB: Provable robustness against backdoor attacks
Maurice Weber, Xiaojun Xu, Bojan Karlaš, Ce Zhang, and Bo Li · 2023
Later among the works it cites.
BadChain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li · 2023
Later among the works it cites.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models, 2023
Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen · 2023
Later among the works it cites.
A comprehensive overview of backdoor attacks in large language models within communication networks, 2023
Haomiao Yang, Kunlan Xiang, Mengyu Ge, Hongwei Li, Rongxing Lu, and Shui Yu · 2023
Later among the works it cites.