Fetching the paper…
Reading the bibliography…
Large language models (LLMs), known for their capability in understanding and following instructions, are vulnerable to adversarial attacks.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
On Randomization in MTD Systems. In Proceedings of the 9th ACM Workshop on Moving Target Defense . 37–43
Majid Ghaderi, Samuel Jero, Cristina Nita-Rotaru, and Reihaneh Safavi-Naini. 2022 · 2022
Earlier work this paper cites.
Hardware Moving Target Defenses against Physical Attacks: Design Challenges and Opportunities. In Proceedings of the 9th ACM Workshop on Moving Target Defense
David S Koblah, Fatemeh Ganji, Domenic Forte, and Shahin Tajik. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Understanding Multi-Turn Toxic Behaviors in Open-Domain Chatbots
Bocheng Chen, Guangjing Wang, Hanqing Guo, Yuanda Wang, and Qiben Yan. 2023 · 2023
Cited alongside, same era.
LLM Guard - The Security Toolkit for LLM Interactions
Laiyer.ai. 2023 · 2023
Cited alongside, same era.
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Closest in time.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…