Fetching the paper…
Reading the bibliography…
Large Language Models (LLMS) have increasingly become central to generating content with potential societal impacts.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
Defending against alignment-breaking attacks via robustly aligned llm
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023 · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023 · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. 2023 · 2023
Earlier work this paper cites.
Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak
Yanrui Du, Sendong Zhao, Ming Ma, Yuhan Chen, and Bing Qin. 2023 · 2023
Earlier work this paper cites.
Llm self defense: By self examination, llms know they are being tricked
Alec Helbling, Mansi Phute, Matthew Hull, and Duen Horng Chau. 2023 · 2023
Earlier work this paper cites.
Token-level adversarial prompt detection based on perplexity measures and contextual information
Zhengmian Hu, Gang Wu, Saayan Mitra, Ruiyi Zhang, Tong Sun, Heng Huang, and Vishy Swaminathan. 2023 · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Earlier work this paper cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. 2023 · 2023
Earlier work this paper cites.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2023 · 2023
Earlier work this paper cites.
Open sesame! universal black box jailbreaking of large language models
Raz Lapid, Ron Langberg, and Moshe Sipper. 2023 · 2023
Earlier work this paper cites.
Vicuna 7b v1.5: A chat assistant fine-tuned on sharegpt conversations
LMSYS. 2023 · 2023
Earlier work this paper cites.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. 2023 · 2023
Earlier work this paper cites.
Lever: Learning to verify language-to-code generation with execution
Ansong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov, Wen-tau Yih, Sida Wang, and Xi Victoria Lin. 2023 · 2023
Cited alongside, same era.
OWASP Top 10 for LLM Applications
OWASP. 2023 · 2023
Cited alongside, same era.
Jatmo: Prompt injection defense by task-specific finetuning
Julien Piet, Maha Alrashed, Chawin Sitawarin, Sizhe Chen, Zeming Wei, Elizabeth Sun, Basel Alomair, and David Wagner. 2023 · 2023
Cited alongside, same era.
Bergeron: Combating adversarial attacks through a conscience-based alignment framework
Matthew Pisano, Peter Ly, Abraham Sanders, Bingsheng Yao, Dakuo Wang, Tomek Strzalkowski, and Mei Si. 2023 · 2023
Cited alongside, same era.
Hijacking large language models via adversarial in-context learning
Yao Qiang, Xiangyu Zhou, and Dongxiao Zhu. 2023 · 2023
Defending large language models against jailbreaking attacks through goal prioritization
Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
FT-Roberta-LLM: A Fine-Tuned Roberta Large Language Model
X. fine tuned. 2024 · 2024
Closest in time.
Catastrophic jailbreak of open-source LLMs via exploiting generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. 2024 · 2024
Closest in time.
Meta llama
Hugging Face. 2023a · 2024
Closest in time.
Vicuna 7b v1.5
Hugging Face. 2023b · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas. 2023 · 2023
Cited alongside, same era.
Adversarial attacks and defenses in large language models: Old and new threats
Leo Schwinn, David Dobre, Stephan Günnemann, and Gauthier Gidel. 2023 · 2023
Cited alongside, same era.
Muhammad Ahmed Shah, Roshan Sharma, Hira Dhamyal, Raphael Olivier, Ankit Shah, Dareen Alharthi, Hazim T Bukhari, Massa Baali, Soham Deshmukh, Michael Kuhlmann, et al. 2023 · 2023
Cited alongside, same era.
Preference ranking optimization for human alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2023 · 2023
Cited alongside, same era.
Why do universal adversarial attacks work on large language models?: Geometry might be the answer
Varshini Subhash, Anna Bialas, Weiwei Pan, and Finale Doshi-Velez. 2023 · 2023
Cited alongside, same era.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Nan Xu, Fei Wang, Ben Zhou, Bang Zheng Li, Chaowei Xiao, and Muhao Chen. 2023 · 2023
Cited alongside, same era.
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. 2024 · 2024
Closest in time.
Moderation guide
OpenAI. 2023 · 2024
Closest in time.
Openai pricing
OpenAI. 2023a · 2024
Closest in time.
Research overview
OpenAI. 2023b · 2024
Closest in time.
Llm-guard
ProtectAI. 2023 · 2024
Closest in time.
Opportunities and challenges for chatgpt and large language models in biomedicine and health
Shubo Tian, Qiao Jin, Lana Yeganova, Po-Ting Lai, Qingqing Zhu, Xiuying Chen, Yifan Yang, Qingyu Chen, Won Kim, Donald C Comeau, et al. 2024 · 2024
Closest in time.
Easyjailbreak: A unified framework for jailbreaking large language models
Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia, Yingshuang Gu, Mingxu Chai, Fukang Zhu, Caishuang Huang, Shihan Dou, Zhiheng Xi, et al. 2024 · 2024
Closest in time.