Fetching the paper…
Reading the bibliography…
Despite extensive pre-training in moral alignment to prevent generating harmful information, large language models (LLMs) remain vulnerable to jailbreak attacks.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston · 2019
Earlier work this paper cites.
Exploring social bias in chatbots using stereotype knowledge
Nayeon Lee, Andrea Madotto, and Pascale Fung · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Anticipating safety issues in e2e conversational ai: Framework and tooling
Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Decomposed prompting: A modular approach for solving complex tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
Defending against alignment-breaking attacks via robustly aligned llm
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch · 2023
Earlier work this paper cites.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al · 2023
Earlier work this paper cites.
Llm self defense: By self examination, llms know they are being tricked
Alec Helbling, Mansi Phute, Matthew Hull, and Duen Horng Chau · 2023
Cited alongside, same era.
Metagpt: Meta programming for multi-agent collaborative framework
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, et al · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang · 2023
Later among the works it cites.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu · 2023
Later among the works it cites.
Defending large language models against jailbreaking attacks through goal prioritization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automatically auditing large language models via discrete optimization
Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi · 2023
Cited alongside, same era.
Black box adversarial prompting for foundation models
Natalie Maus, Patrick Chao, Eric Wong, and Jacob R Gardner · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2023
Cited alongside, same era.
Gpt-4 technical report
R OpenAI · 2023
Cited alongside, same era.
Bergeron: Combating adversarial attacks through a conscience-based alignment framework
Matthew Pisano, Peter Ly, Abraham Sanders, Bingsheng Yao, Dakuo Wang, Tomek Strzalkowski, and Mei Si · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Cited alongside, same era.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas · 2023
Cited alongside, same era.
Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
The impact of reasoning step length on large language models
Mingyu Jin, Qinkai Yu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, Mengnan Du, et al · 2024
Closest in time.
Protecting your llms with information bottleneck
Zichuan Liu, Zefan Wang, Linjie Xu, Jinyu Wang, Lei Song, Tianchun Wang, Chunlin Chen, Wei Cheng, and Jiang Bian · 2024
Closest in time.
Advprompter: Fast adaptive adversarial prompting for llms
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Closest in time.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel · 2024
Closest in time.
Defending llms against jailbreaking attacks via backtranslation
Yihan Wang, Zhouxing Shi, Andrew Bai, and Cho-Jui Hsieh · 2024
Closest in time.
Llm jailbreak attack versus defense techniques–a comprehensive study
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek · 2024
Closest in time.