Fetching the paper…
Reading the bibliography…
The aligned Large Language Models (LLMs) are powerful language understanding and decision-making tools that are created through extensive alignment with human feedback.
Transferability in machine learning: from phenomena to black-box attacks using adversarial samples
Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow · 2016
Earlier work this paper cites.
Finbert: Financial sentiment analysis with pre-trained language models
Dogu Araci · 2019
Earlier work this paper cites.
Reevaluating adversarial examples in natural language
John X Morris, Eli Lifland, Jack Lanchantin, Yangfeng Ji, and Yanjun Qi · 2020
Earlier work this paper cites.
Recursively summarizing books with human feedback
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Earlier work this paper cites.
Promptsource: An integrated development environment and repository for natural language prompts, 2022
Stephen H. Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Alan Fries, Maged S. Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Xiangru Tang, Mike Tian-Jian Jiang, and Alexander M. Rush · 2022
Earlier work this paper cites.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
Biogpt: generative pre-trained transformer for biomedical text generation and mining
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
https://www.jailbreakchat.com/ , 2023
Alex Albert · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity, 2023
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
The hacking of chatgpt is just getting started
Matt Burgess · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Cited alongside, same era.
Amazing “jailbreak” bypasses chatgpt’s ethics safeguards
Jon Christian · 2023
Cited alongside, same era.
Jailbreaker: Automated jailbreak across multiple large language model chatbots, 2023
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2023
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Cited alongside, same era.
Jailbreaking chatgpt via prompt engineering: An empirical study, 2023
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu · 2023
Closest in time.
https://gist.github.com/coolaj86/6f4f7b30129b0251f61fa7baaa881516 , 2023
AJ ONeal · 2023
Closest in time.
Snapshot of gpt-3.5-turbo from march 1st 2023
OpenAI · 2023
Closest in time.
Fine-tuning large neural language models for biomedical natural language processing
Robert Tinn, Hao Cheng, Yu Gu, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Josh A Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova · 2023
Cited alongside, same era.
https://huggingface.co/datasets/Dahoas/synthetic-instruct-gptj-pairwise , 2023
Alex Havrilla · 2023
Cited alongside, same era.
Large language models can be used to effectively scale spear phishing campaigns
Julian Hazell · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models, 2023
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2023
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models, 2023
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Cited alongside, same era.
Multi-step jailbreaking privacy attacks on chatgpt, 2023
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song · 2023
Cited alongside, same era.
https://old.reddit.com/r/ChatGPT/comments/zlcyr9/dan_is_my_new_friend/ , 2022
walkerspider · 2023
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2023
Closest in time.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
Easyjailbreak: A unified framework for jailbreaking large language models
Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia, Yingshuang Gu, Mingxu Chai, Fukang Zhu, Caishuang Huang, Shihan Dou, Zhiheng Xi, Rui Zheng, Songyang Gao, Yicheng Zou, Hang Yan, Yifan Le, Ruohui Wang, Lijun Li, Jing Shao, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Closest in time.