Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) display versatile functionality, they continue to generate harmful, biased, and toxic content, as demonstrated by the prevalence of human-designed jailbreaks.
“Fine-Tuning Language Models from Human Preferences”, 2020
Daniel. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom. Brown, Alec Radford, Dario Amodei, Paul Christiano and Geoffrey Irving · 1909
Earlier work this paper cites.
“Fine-Tuning Language Models from Human Preferences”, 2020
Daniel. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom. Brown, Alec Radford, Dario Amodei, Paul Christiano and Geoffrey Irving · 1909
Earlier work this paper cites.
“HotFlip: White-Box Adversarial Examples for Text Classification”
Javid Ebrahimi, Anyi Rao, Daniel Lowd and Dejing Dou · 2006
Earlier work this paper cites.
“HotFlip: White-Box Adversarial Examples for Text Classification”
Javid Ebrahimi, Anyi Rao, Daniel Lowd and Dejing Dou · 2006
Earlier work this paper cites.
“Recipes for Safety in Open-domain Chatbots”, 2021
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston and Emily Dinan · 2010
Earlier work this paper cites.
“Recipes for Safety in Open-domain Chatbots”, 2021
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston and Emily Dinan · 2010
Earlier work this paper cites.
“Intriguing Properties of Neural Networks”
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian. Goodfellow and Rob Fergus · 2014
Earlier work this paper cites.
“Intriguing Properties of Neural Networks”
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian. Goodfellow and Rob Fergus · 2014
Earlier work this paper cites.
“Explaining and Harnessing Adversarial Examples”
Ian. Goodfellow, Jonathon Shlens and Christian Szegedy · 2015
Earlier work this paper cites.
“Explaining and Harnessing Adversarial Examples”
Ian. Goodfellow, Jonathon Shlens and Christian Szegedy · 2015
Earlier work this paper cites.
“Concrete Problems in AI Safety”, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman and Dan Mané · 2016
Earlier work this paper cites.
“Concrete Problems in AI Safety”, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman and Dan Mané · 2016
Earlier work this paper cites.
“Deep Reinforcement Learning from Human Preferences”
Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg and Dario Amodei · 2017
Earlier work this paper cites.
“Deep Reinforcement Learning from Human Preferences”
Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg and Dario Amodei · 2017
Earlier work this paper cites.
“Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers”
Ji Gao, Jack Lanchantin, Mary Soffa and Yanjun Qi · 2018
Earlier work this paper cites.
“Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers”
Ji Gao, Jack Lanchantin, Mary Soffa and Yanjun Qi · 2018
Earlier work this paper cites.
“TextBugger: Generating Adversarial Text Against Real-world Applications”
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li and Ting Wang · 2019
Earlier work this paper cites.
“Combating Adversarial Misspellings with Robust Word Recognition”
Danish Pruthi, Bhuwan Dhingra and Zachary. Lipton · 2019
Earlier work this paper cites.
“The Woman Worked as a Babysitter: On Biases in Language Generation”
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan and Nanyun Peng · 2019
Earlier work this paper cites.
“Universal Adversarial Triggers for Attacking and Analyzing NLP”
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner and Sameer Singh · 2019
Earlier work this paper cites.
“TextBugger: Generating Adversarial Text Against Real-world Applications”
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li and Ting Wang · 2019
Earlier work this paper cites.
“Combating Adversarial Misspellings with Robust Word Recognition”
Danish Pruthi, Bhuwan Dhingra and Zachary. Lipton · 2019
Earlier work this paper cites.
“The Woman Worked as a Babysitter: On Biases in Language Generation”
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan and Nanyun Peng · 2019
Earlier work this paper cites.
“Universal Adversarial Triggers for Attacking and Analyzing NLP”
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner and Sameer Singh · 2019
Earlier work this paper cites.
“Language Models are Few-Shot Learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever and Dario Amodei · 2020
Earlier work this paper cites.
“Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel and Douwe Kiela · 2020
Earlier work this paper cites.
“BERT-ATTACK: Adversarial Attack Against BERT Using BERT”
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue and Xipeng Qiu · 2020
Earlier work this paper cites.
“It’s Morphin’ Time! Combating Linguistic Discrimination with Inflectional Perturbations”
Samson Tan, Shafiq Joty, Min-Yen Kan and Richard Socher · 2020
Earlier work this paper cites.
“Word-level Textual Adversarial Attacking as Combinatorial Optimization”
Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu and Maosong Sun · 2020
Earlier work this paper cites.
“Language Models are Few-Shot Learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever and Dario Amodei · 2020
Earlier work this paper cites.
“BERT-ATTACK: Adversarial Attack Against BERT Using BERT”
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue and Xipeng Qiu · 2020
Earlier work this paper cites.
“Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel and Douwe Kiela · 2020
Earlier work this paper cites.
“It’s Morphin’ Time! Combating Linguistic Discrimination with Inflectional Perturbations”
Samson Tan, Shafiq Joty, Min-Yen Kan and Richard Socher · 2020
Earlier work this paper cites.
“Word-level Textual Adversarial Attacking as Combinatorial Optimization”
Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu and Maosong Sun · 2020
Earlier work this paper cites.
“Persistent Anti-Muslim Bias in Large Language Models”
Abubakar Abid, Maheen Farooqi and James Zou · 2021
Earlier work this paper cites.
“On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”
Emily. Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
“Extracting Training Data from Large Language Models”
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea and Colin Raffel · 2021
Earlier work this paper cites.
“Gradient-based Adversarial Attacks against Text Transformers”
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou and Douwe Kiela · 2021
Earlier work this paper cites.
“Aligning AI With Shared Human Values”
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song and Jacob Steinhardt · 2021
Earlier work this paper cites.
“Universal Adversarial Attacks with Natural Triggers for Text Classification”
Liwei Song, Xinwei Yu, Hsuan-Tung Peng and Karthik Narasimhan · 2021
Earlier work this paper cites.
“GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model”, 2021
Ben Wang and Aran Komatsuzaki · 2021
Cited alongside, same era.
“Persistent Anti-Muslim Bias in Large Language Models”
Abubakar Abid, Maheen Farooqi and James Zou · 2021
Cited alongside, same era.
“On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”
Emily. Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell · 2021
Cited alongside, same era.
“Extracting Training Data from Large Language Models”
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea and Colin Raffel · 2021
Cited alongside, same era.
“Gradient-based Adversarial Attacks against Text Transformers”
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou and Douwe Kiela · 2021
Cited alongside, same era.
“Aligning AI With Shared Human Values”
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song and Jacob Steinhardt · 2021
“Universal and Transferable Adversarial Attacks on Aligned Language Models”, 2023
Andy Zou, Zifan Wang, J. Kolter and Matt Fredrikson · 2023
Closest in time.
“Detecting Language Model Attacks with Perplexity”, 2023
Gabriel Alon and Michael Kamfonas · 2023
Closest in time.
“Jailbreak Chat”, 2023
Jailbreak Chat · 2023
Closest in time.
“Explore, Establish, Exploit: Red Teaming Language Models from Scratch”, 2023
Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp and Dylan Hadfield-Menell · 2023
Closest in time.
“Jailbreaking Black Box Large Language Models in Twenty Queries”, 2023
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George. Pappas and Eric Wong · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Universal Adversarial Attacks with Natural Triggers for Text Classification”
Liwei Song, Xinwei Yu, Hsuan-Tung Peng and Karthik Narasimhan · 2021
Cited alongside, same era.
“GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model”, 2021
Ben Wang and Aran Komatsuzaki · 2021
Cited alongside, same era.
“Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback”, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann and Jared Kaplan · 2022
Cited alongside, same era.
“Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned”, 2022
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan and Jack Clark · 2022
Cited alongside, same era.
“Debiased Large Language Models Still Associate Muslims with Uniquely Violent Acts”, 2022
Babak Hemmatian and Lav. Varshney · 2022
Cited alongside, same era.
“DAN is my new friend”, 2022
walkerspider · 2022
Cited alongside, same era.
Closest in time.
“Toxicity in ChatGPT: Analyzing Persona-Assigned Language Models”
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan and Karthik Narasimhan · 2023
Closest in time.
“Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations”, 2023
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine and Madian Khabsa · 2023
Closest in time.
“Automatically Auditing Large Language Models via Discrete Optimization”
Erik Jones, Anca Dragan, Aditi Raghunathan and Jacob Steinhardt · 2023
Closest in time.
“BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset”
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang and Yaodong Yang · 2023
Closest in time.
“Gender Bias and Stereotypes in Large Language Models”
Hadas Kotek, Rikker Dockum and David Sun · 2023
Closest in time.
“Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study”, 2023
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang and Yang Liu · 2023
Closest in time.
“Multi-step Jailbreaking Privacy Attacks on ChatGPT”
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng and Yangqiu Song · 2023
Closest in time.
“Open Sesame! Universal Black Box Jailbreaking of Large Language Models”, 2023
Raz Lapid, Ron Langberg and Moshe Sipper · 2023
Closest in time.
“Hijacking Large Language Models via Adversarial In-Context Learning”, 2023
Yao Qiang, Xiangyu Zhou and Dongxiao Zhu · 2023
Closest in time.
“Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation”, 2023
Rusheb Shah, Quentin Feuillade–Montixi, Soroush Pour, Arush Tagade, Stephen Casper and Javier Rando · 2023
Closest in time.
Muhammad Shah, Roshan Sharma, Hira Dhamyal, Raphael Olivier, Ankit Shah, Joseph Konan, Dareen Alharthi, Hazim Bukhari, Massa Baali, Soham Deshmukh, Michael Kuhlmann, Bhiksha Raj and Rita Singh · 2023
Closest in time.
“Llama 2: Open Foundation and Fine-Tuned Chat Models”, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Smith, Ranjan Subramanian, Xiaoqing Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov and Thomas Scialom · 2023
Closest in time.
“Jailbroken: How Does LLM Safety Training Fail?”
Alexander Wei, Nika Haghtalab and Jacob Steinhardt · 2023
Closest in time.
“You can use GPT-4 to create prompt injections against GPT-4”, 2023
WitchBOT · 2023
Closest in time.
“Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery”
Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping and Tom Goldstein · 2023
Closest in time.
“Aligning Large Language Models with Human: A Survey”, 2023
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang and Qun Liu · 2023
Closest in time.
“GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts”, 2023
Jiahao Yu, Xingwei Lin, Zheng Yu and Xinyu Xing · 2023
Closest in time.
“Tree of Thoughts: Deliberate Problem Solving with Large Language Models”
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao and Karthik Narasimhan · 2023
Closest in time.
“Universal and Transferable Adversarial Attacks on Aligned Language Models”, 2023
Andy Zou, Zifan Wang, J. Kolter and Matt Fredrikson · 2023
Closest in time.
“Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks”, 2024
Maksym Andriushchenko, Francesco Croce and Nicolas Flammarion · 2024
Closest in time.
“MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots”
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang and Yang Liu · 2024
Closest in time.
“Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks”
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia and Tatsunori Hashimoto · 2024
Closest in time.
“AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models”
Xiaogeng Liu, Nan Xu, Muhao Chen and Chaowei Xiao · 2024
Closest in time.
“HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal”
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth and Dan Hendrycks · 2024
Closest in time.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen and Yang Zhang · 2024
Closest in time.
“A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models”
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li and Stjepan Picek · 2024
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia and Weiyan Shi · 2024
Closest in time.
“Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks”, 2024
Maksym Andriushchenko, Francesco Croce and Nicolas Flammarion · 2024
Closest in time.
“MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots”
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang and Yang Liu · 2024
Closest in time.
“Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks”
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia and Tatsunori Hashimoto · 2024
Closest in time.
“AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models”
Xiaogeng Liu, Nan Xu, Muhao Chen and Chaowei Xiao · 2024
Closest in time.
“HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal”
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth and Dan Hendrycks · 2024
Closest in time.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen and Yang Zhang · 2024
Closest in time.
“A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models”
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li and Stjepan Picek · 2024
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia and Weiyan Shi · 2024
Closest in time.