Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have been shown to be susceptible to jailbreak attacks, or adversarial attacks used to illicit high risk behavior from a model.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu. 2019 · 1907
Earlier work this paper cites.
Situating sentence embedders with nearest neighbor overlap
Lucy H Lin and Noah A Smith. 2019 · 1909
Earlier work this paper cites.
Active learning with statistical models
David A Cohn, Zoubin Ghahramani, and Michael I Jordan. 1996 · 1996
Earlier work this paper cites.
Gedi: Generative discriminator guided sequence generation
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2020 · 2009
Earlier work this paper cites.
A systematic review and meta-analysis of the effectiveness of nudging to increase fruit and vegetable choice
Valérie JV Broers, Céline De Breucker, Stephan Van den Broucke, and Olivier Luminet. 2017 · 2017
Earlier work this paper cites.
Machine teaching: A new paradigm for building machine learning systems
Patrice Y Simard, Saleema Amershi, David M Chickering, Alicia Edelman Pelton, Soroush Ghorashi, Christopher Meek, Gonzalo Ramos, Jina Suh, Johan Verwey, Mo Wang, et al. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani. 2017 · 2017
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Preventing repeated real world ai failures by cataloging incidents: The ai incident database
Sean McGregor. 2021 · 2021
Earlier work this paper cites.
Nudging toward vaccination: a systematic review
Mark Donald C Reñosa, Jeniffer Landicho, Jonas Wachinger, Sarah L Dalglish, Kate Bärnighausen, Till Bärnighausen, and Shannon A McMahon. 2021 · 2021
Earlier work this paper cites.
Fudge: Controlled text generation with future discriminators
Kevin Yang and Dan Klein. 2021 · 2021
Earlier work this paper cites.
Critic-guided decoding for controlled text generation
Minbeom Kim, Hwanhee Lee, Kang Min Yoo, Joonsuk Park, Hwaran Lee, and Kyomin Jung. 2022 · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Cited alongside, same era.
Controllable natural language generation with contrastive prefixes
Jing Qian, Li Dong, Yelong Shen, Furu Wei, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Out-of-distribution detection and selective generation for conditional language models
Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu. 2022 · 2022
Cited alongside, same era.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas. 2023 · 2023
Cited alongside, same era.
Llm censorship: A machine learning challenge or a computer security problem?
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2023 · 2023
Later among the works it cites.
Instruction-following evaluation for large language models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
How many opinions does your llm have? improving uncertainty estimation in nlg
Lukas Aichberger, Kajetan Schweighofer, Mykyta Ielanskyi, and Sepp Hochreiter. 2024 · 2024
Later among the works it cites.
Foundational challenges in assuring alignment and safety of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Glukhov, Ilia Shumailov, Yarin Gal, Nicolas Papernot, and Vardan Papyan. 2023 · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. 2023 · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2023 · 2023
Cited alongside, same era.
Bergeron: Combating adversarial attacks through a conscience-based alignment framework
Matthew Pisano, Peter Ly, Abraham Sanders, Bingsheng Yao, Dakuo Wang, Tomek Strzalkowski, and Mei Si. 2023 · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023 · 2023
Cited alongside, same era.
Self-guard: Empower the llm to safeguard itself
Zezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wang, Liang Chen, Qingwei Lin, and Kam-Fai Wong. 2023 · 2023
Cited alongside, same era.
Jailbreak and guard aligned language models with only few in-context demonstrations
Zeming Wei, Yifei Wang, Ang Li, Yichuan Mo, and Yisen Wang. 2023 · 2023
Cited alongside, same era.
A framework for real-time safeguarding the text generation of large language
Ximing Dong, Dayi Lin, Shaowei Wang, and Ahmed E Hassan. 2024a
Cited in the paper.
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al. 2024 · 2024
Later among the works it cites.
Nudging: Inference-time alignment via model collaboration
Yu Fei, Yasaman Razeghi, and Sameer Singh. 2024 · 2024
Later among the works it cites.
Infusing behavior science into large language models for activity coaching
Narayan Hegde, Madhurima Vardhan, Deepak Nathani, Emily Rosenzweig, Cathy Speed, Alan Karthikesalingam, and Martin Seneviratne. 2024 · 2024
Later among the works it cites.
Malla: Demystifying real-world large language model integrated malicious services
Zilong Lin, Jian Cui, Xiaojing Liao, and XiaoFeng Wang. 2024 · 2024
Later among the works it cites.
Cbf-llm: Safe control for llm alignment
Yuya Miyaoka and Masaki Inoue. 2024 · 2024
Later among the works it cites.
Guobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He, and Yi Zeng. 2024 · 2024
Later among the works it cites.
A comprehensive study of jailbreak attack versus defense for large language models
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek. 2024 · 2024
Later among the works it cites.