Fetching the paper…
Reading the bibliography…
Backdoors are hidden behaviors that are only triggered once an AI system has been deployed.
Backdoor attacks against learning systems
Yujie Ji, Xinyang Zhang, and Ting Wang · 2017
Earlier work this paper cites.
Concealed data poisoning attacks on NLP models
Eric Wallace, Tony Z Zhao, Shi Feng, and Sameer Singh · 2020
Earlier work this paper cites.
Risks from Learned Optimization in Advanced Machine Learning Systems, 2021
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2021
Earlier work this paper cites.
Hidden backdoors in human-centric language models, 2021
Shaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao, Minhui Xue, Haojin Zhu, and Jialiang Lu · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in LLMs, 2023
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans · 2023
Earlier work this paper cites.
Poisoning web-scale training datasets is practical
Nicholas Carlini, Matthew Jagielski, Christopher A Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr · 2023
Earlier work this paper cites.
Composite backdoor attacks against large language models
Hai Huang, Zhengyu Zhao, Michael Backes, Yun Shen, and Yang Zhang · 2023
Earlier work this paper cites.
Model organisms of misalignment: The case for a new pillar of alignment research
Evan Hubinger, Nicholas Schiefer, Carson Denison, and Ethan Perez · 2023
Cited alongside, same era.
Time is Encoded in the Weights of Finetuned Language Models, 2023
Kai Nylund, Suchin Gururangan, and Noah A. Smith · 2023
Cited alongside, same era.
Instruction Tuning with GPT-4, 2023
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Cited alongside, same era.
Universal jailbreak backdoors from poisoned human feedback
Javier Rando and Florian Tramèr · 2023
Cited alongside, same era.
OpenHermes 2.5: An Open Dataset of Synthetic Data for Generalist LLM Assistants, 2023
Teknium · 2023
Cited alongside, same era.
Activation Addition: Steering Language Models Without Optimization, 2023
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning, 2024
Bahare Fatemi, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan, Jinyeong Yim, John Palowitch, Sungyong Seo, Jonathan Halcrow, and Bryan Perozzi · 2024
Closest in time.
Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration, 2024
Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Vidhisha Balachandran, and Yulia Tsvetkov · 2024
Closest in time.
Language Models Represent Space and Time, 2024
Wes Gurnee and Max Tegmark · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training, 2024
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
Steering Llama 2 via Contrastive Activation Addition, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Cited alongside, same era.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Cited alongside, same era.
Do Large Language Models Know What They Don’t Know?, 2023
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang · 2023
Cited alongside, same era.
Dated Data: Tracing Knowledge Cutoffs in Large Language Models, 2024
Jeffrey Cheng, Marc Marone, Orion Weller, Dawn Lawrie, Daniel Khashabi, and Benjamin Van Durme · 2024
Cited alongside, same era.
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.
Latent adversarial training improves robustness to persistent harmful behaviors in llms, 2024
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, and Stephen Casper · 2024
Closest in time.
New york times archive
The New York Times · 2024
Closest in time.
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun · 2024
Closest in time.
Rapid adoption, hidden risks: The dual impact of large language model customization
Rui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang, Yuan Zhang, Michael Backes, Yun Shen, and Yang Zhang · 2024
Closest in time.