Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly deployed in various applications.
“Perplexity—a measure of the difficulty of speech recognition tasks”
Fred Jelinek, Robert Mercer, Lalit Bahl and James Baker · 1977
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
Zhuang Ma and Michael Collins · 2018
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models”
Edward Hu et al · 2021
Earlier work this paper cites.
“A survey on in-context learning”
Qingxiu Dong et al · 2022
Earlier work this paper cites.
“Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned”
Deep Ganguli et al · 2022
Earlier work this paper cites.
Thomas Hartvigsen et al · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Earlier work this paper cites.
“Red teaming language models with language models”
Ethan Perez et al · 2022
Earlier work this paper cites.
“Falcon-40B: an open large language model with state-of-the-art performance”, 2023
Ebtesam Almazrouei et al · 2023
Earlier work this paper cites.
“Baichuan 2: Open Large-scale Language Models”
Baichuan · 2023
Earlier work this paper cites.
“Explore, establish, exploit: Red teaming language models from scratch”
Stephen Casper et al · 2023
Earlier work this paper cites.
“Jailbreaking black box large language models in twenty queries”
Patrick Chao et al · 2023
Earlier work this paper cites.
“Attack prompt generation for red teaming and defending large language models”
Boyi Deng et al · 2023
Earlier work this paper cites.
“Jailbreaker: Automated jailbreak across multiple large language model chatbots”
Gelei Deng et al · 2023
Earlier work this paper cites.
“Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak”
Yanrui Du et al · 2023
Cited alongside, same era.
“Llm self defense: By self examination, llms know they are being tricked”
Alec Helbling, Mansi Phute, Matthew Hull and Duen Chau · 2023
Cited alongside, same era.
“Baseline defenses for adversarial attacks against aligned language models”
Neel Jain et al · 2023
Cited alongside, same era.
“Automatically auditing large language models via discrete optimization”
Erik Jones, Anca Dragan, Aditi Raghunathan and Jacob Steinhardt · 2023
Cited alongside, same era.
“Exploiting programmatic behavior of llms: Dual-use through standard security attacks”
“Defending chatgpt against jailbreak attack via self-reminder”, 2023
Fangzhao Wu et al · 2023
Later among the works it cites.
“Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts”
Jiahao Yu, Xingwei Lin and Xinyu Xing · 2023
Later among the works it cites.
“Judging LLM-as-a-judge with MT-Bench and Chatbot Arena”
Lianmin Zheng et al · 2023
Later among the works it cites.
“Universal and transferable adversarial attacks on aligned language models”
Andy Zou, Zifan Wang, J Kolter and Matt Fredrikson · 2023
Later among the works it cites.
“Attacks, defenses and evaluations for llm conversation safety: A survey”
Zhichen Dong et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Kang et al · 2023
Cited alongside, same era.
“Certifying llm safety against adversarial prompting”
Aounon Kumar et al · 2023
Cited alongside, same era.
“The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning”
Bill Lin et al · 2023
Cited alongside, same era.
“Autodan: Generating stealthy jailbreak prompts on aligned large language models”
Xiaogeng Liu, Nan Xu, Muhao Chen and Chaowei Xiao · 2023
Cited alongside, same era.
“A holistic approach to undesired content detection in the real world”
Todor Markov et al · 2023
Cited alongside, same era.
“Flirt: Feedback loop in-context red teaming”
Ninareh Mehrabi et al · 2023
Cited alongside, same era.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Cited alongside, same era.
“Fine-tuning aligned language models compromises safety, even when users do not intend to!”
Xiangyu Qi et al · 2023
Cited alongside, same era.
Closest in time.
“Cold-attack: Jailbreaking llms with stealthiness and controllability”
Xingang Guo et al · 2024
Closest in time.
Albert Jiang et al · 2024
Closest in time.
“Tuning language models by proxy”
Alisa Liu et al · 2024
Closest in time.
“Warm: On the benefits of weight averaged reward models”
Alexandre Ramé et al · 2024
Closest in time.
“Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology”
Zhenhua Wang et al · 2024
Closest in time.
“Jailbroken: How does llm safety training fail?”
Alexander Wei, Nika Haghtalab and Jacob Steinhardt · 2024
Closest in time.
“SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding”
Zhangchen Xu et al · 2024
Closest in time.
“LLM Jailbreak Attack versus Defense Techniques–A Comprehensive Study”
Zihao Xu et al · 2024
Closest in time.
“Weak-to-strong jailbreaking on large language models”
Xuandong Zhao et al · 2024
Closest in time.
“Lima: Less is more for alignment”
Chunting Zhou et al · 2024
Closest in time.