Fetching the paper…
Reading the bibliography…
Generative models are rapidly gaining popularity and being integrated into everyday applications, raising concerns over their safe use as various vulnerabilities are exposed.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
Svetlana Kiritchenko and Saif M. Mohammad · 2005
Earlier work this paper cites.
Computational red teaming: Past, present and future, 2011
Hussein Abbass, Axel Bender, Svetoslav Gaidow, and Paul Whitbread · 2011
Earlier work this paper cites.
Mapping the moral domain
Jesse Graham, Brian A Nosek, Jonathan Haidt, Ravi Iyer, Spassena Koleva, and Peter H Ditto · 2011
Earlier work this paper cites.
Sequence to sequence learning with neural networks, 2014
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2016
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks, 2017
Aleksander Madry · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Deep feature interpolation for image content changes
Paul Upchurch, Jacob Gardner, Geoff Pleiss, Robert Pless, Noah Snavely, Kavita Bala, and Kilian Weinberger · 2017
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Adversarial attacks and defenses in deep learning
Kui Ren, Tianhang Zheng, Zhan Qin, and Xue Liu · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment, 2021
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Earlier work this paper cites.
HateBERT: Retraining BERT for abusive language detection in English
Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, and Clemens Winter · 2021
Earlier work this paper cites.
Diverse adversaries for mitigating bias in training
Xudong Han, Timothy Baldwin, and Trevor Cohn · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Earlier work this paper cites.
Bot-adversarial dialogue for safe conversational agents
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan · 2021
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
Capturing failures of large language models via human cognitive biases, 2022
Erik Jones and Jacob Steinhardt · 2022
Earlier work this paper cites.
Researching alignment research: Unsupervised analysis
Jan H. Kirchner, Logan Smith, Jacques Thibodeau, Kyle McDonell, and Laria Reynolds · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Fast model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2022
Earlier work this paper cites.
A survey of machine unlearning, 2022
Thanh Tam Nguyen, Thanh Trung Huynh, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, H. Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
Learning to retrieve prompts for in-context learning
Ohad Rubin, Jonathan Herzig, and Jonathan Berant · 2022
Earlier work this paper cites.
Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques
Mohammed Latif Siddiq and Joanna CS Santos · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Earlier work this paper cites.
Physics of language models: Part 3.2, knowledge manipulation
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, et al · 2023
Earlier work this paper cites.
DICES dataset: Diversity in conversational AI evaluation for safety
Lora Aroyo, Alex S. Taylor, Mark Díaz, Christopher Homan, Alicia Parrish, Gregory Serapio-García, Vinodkumar Prabhakaran, and Ding Wang · 2023
Earlier work this paper cites.
Abusing images and sounds for indirect instruction injection in multi-modal llms, 2023
Eugene Bagdasaryan, Tsung-Yin Hsieh, Ben Nassi, and Vitaly Shmatikov · 2023
Earlier work this paper cites.
The reversal curse: Llms trained on ”a is b” fail to learn ”b is a”
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2023
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models, 2023
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, Sasha Frolov, Ravi Prakash Giri, Dhaval Kapil, Yiannis Kozyrakis, David LeBlanc, James Milazzo, Aleksandar Straumann, Gabriel Synnaeve, Varun Vontimitta, Spencer Whitman, and Joshua Saxe · 2023
Earlier work this paper cites.
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou · 2023
Earlier work this paper cites.
Deceptive alignment monitoring
Andres Carranza, Dhruv Pai, Rylan Schaeffer, Arnuv Tandon, and Sanmi Koyejo · 2023
Earlier work this paper cites.
Explore, establish, exploit: Red teaming language models from scratch
Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, and Dylan Hadfield-Menell · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Jailbreaker in jail: Moving target defense for large language models
Bocheng Chen, Advait Paliwal, and Qiben Yan · 2023
Earlier work this paper cites.
Prompting4debugging: Red-teaming text-to-image diffusion models by finding problematic prompts
Zhi-Yi Chin, Chieh-Ming Jiang, Ching-Chun Huang, Pin-Yu Chen, and Wei-Chen Chiu · 2023
Earlier work this paper cites.
Free dolly: Introducing the world’s first truly open instruction-tuned llm
Mike Conover, Matt Hayes, Ankit Mathur, Xiangrui Meng, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin · 2023
Earlier work this paper cites.
Added toxicity mitigation at inference time for multimodal and massively multilingual translation
Marta R. Costa-jussà, David Dale, Maha Elbayad, and Bokai Yu · 2023
Earlier work this paper cites.
On the robustness of large multimodal models against image adversarial attacks
Xuanming Cui, Alejandro Aparcedo, Young Kyun Jang, and Ser-Nam Lim · 2023
Earlier work this paper cites.
Safe RLHF: safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2023
Earlier work this paper cites.
Attack prompt generation for red teaming and defending large language models
Boyi Deng, Wenjie Wang, Fuli Feng, Yang Deng, Qifan Wang, and Xiangnan He · 2023
Earlier work this paper cites.
Assessing language model deployment with risk cards, 2023
Leon Derczynski, Hannah Rose Kirk, Vidhisha Balachandran, Sachin Kumar, Yulia Tsvetkov, Michael Leiser, and Saif Mohammad · 2023
Earlier work this paper cites.
Beyond the safeguards: Exploring the security risks of chatgpt
Erik Derner and Kristina Batistic · 2023
Earlier work this paper cites.
A security risk taxonomy for large language models
Erik Derner, Kristina Batistic, Jan Zahálka, and Robert Babuska · 2023
Earlier work this paper cites.
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang · 2023
Earlier work this paper cites.
How robust is google’s bard to adversarial image attacks?
Yinpeng Dong, Huanran Chen, Jiawei Chen, Zhengwei Fang, Xiao Yang, Yichi Zhang, Yu Tian, Hang Su, and Jun Zhu · 2023
Earlier work this paper cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Earlier work this paper cites.
Aysan Esmradi, Daniel Wankit Yip, and Chun-Fai Chan · 2023
Earlier work this paper cites.
Efficient black-box adversarial attacks on neural text detectors
Vitalii Fishchuk and Daniel Braun · 2023
Earlier work this paper cites.
Scaling laws for adversarial attacks on language model activations
Stanislav Fort · 2023
Earlier work this paper cites.
Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b
Pranav Gade, Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish · 2023
Earlier work this paper cites.
MART: improving LLM safety with multi-round automatic red-teaming
Suyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, and Yuning Mao · 2023
Earlier work this paper cites.
Figstep: Jailbreaking large vision-language models via typographic visual prompts
Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang · 2023
Earlier work this paper cites.
AI control: Improving safety despite intentional subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger · 2023
Earlier work this paper cites.
From chatgpt to threatgpt: Impact of generative AI in cybersecurity and privacy
Maanak Gupta, Charankumar Akiri, Kshitiz Aryal, Eli Parker, and Lopamudra Praharaj · 2023
Earlier work this paper cites.
Large language models for code: Security hardening and adversarial testing
Jingxuan He and Martin Vechev · 2023
Earlier work this paper cites.
An empirical study of metrics to measure representational harms in pre-trained language models
Saghar Hosseini, Hamid Palangi, and Ahmed Hassan Awadallah · 2023
Earlier work this paper cites.
Token-level adversarial prompt detection based on perplexity measures and contextual information
Zhengmian Hu, Gang Wu, Saayan Mitra, Ruiyi Zhang, Tong Sun, Heng Huang, and Vishy Swaminathan · 2023
Earlier work this paper cites.
Robustness tests for automatic machine translation metrics with adversarial attacks
Yichen Huang and Timothy Baldwin · 2023
Earlier work this paper cites.
Walking a tightrope - evaluating large language models in high-risk domains
Chia-Chien Hung, Wiem Ben Rim, Lindsay Frost, Lars Brückner, and Carolin Lawrence · 2023
Earlier work this paper cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Earlier work this paper cites.
Summon a demon and bind it: A grounded theory of LLM red teaming in the wild
Nanna Inie, Jonathan Stray, and Leon Derczynski · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Earlier work this paper cites.
Hijacking context in large multi-modal models
Joonhyun Jeong · 2023
Earlier work this paper cites.
Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Zhang, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Earlier work this paper cites.
Llmlingua: Compressing prompts for accelerated inference of large language models
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu · 2023
Earlier work this paper cites.
Backdoor attacks for in-context learning with language models
Nikhil Kandpal, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini · 2023
Earlier work this paper cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2023
Earlier work this paper cites.
Learn what NOT to learn: Towards generative safety in chatbots
Leila Khalatbari, Yejin Bang, Dan Su, Willy Chung, Saeed Ghadimi, Hossein Sameti, and Pascale Fung · 2023
Earlier work this paper cites.
Robust safety classifier for large language models: Adversarial prompt shield
Jinhwa Kim, Ali Derakhshan, and Ian G. Harris · 2023
Earlier work this paper cites.
Evaluating language-model agents on realistic autonomous tasks
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R. Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, Elizabeth Barnes, and Paul Christiano · 2023
Earlier work this paper cites.
Certifying llm safety against adversarial prompting, 2023
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju · 2023
Cited alongside, same era.
Entangled preferences: The history and risks of reinforcement learning and human feedback
Nathan Lambert, Thomas Krendl Gilbert, and Tom Zick · 2023
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Cited alongside, same era.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish · 2023
Cited alongside, same era.
Multi-step jailbreaking privacy attacks on chatgpt
Interpretability of LLM deception: Universal motif
Anonymous · 2024
Closest in time.
GPT in sheep’s clothing: The risk of customized gpts
Sagiv Antebi, Noam Azulay, Edan Habler, Ben Ganon, Asaf Shabtai, and Yuval Elovici · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Closest in time.
Somnath Banerjee, Sayan Layek, Rima Hazra, and Animesh Mukherjee · 2024
Closest in time.
Detection and defense against prominent attacks on preconditioned llm-integrated virtual assistants
Chun-Fai Chan, Daniel Wankit Yip, and Aysan Esmradi · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song · 2023
Cited alongside, same era.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Cited alongside, same era.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang · 2023
Cited alongside, same era.
Red teaming game: A game-theoretic framework for red teaming language models
Chengdong Ma, Ziran Yang, Minquan Gao, Hai Ci, Jun Gao, Xuehai Pan, and Yaodong Yang · 2023
Cited alongside, same era.
Adversarial prompting for black box foundation models
Natalie Maus, Patrick Chao, Eric Wong, and Jacob R. Gardner · 2023
Cited alongside, same era.
R. Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L. Griffiths · 2023
Cited alongside, same era.
Using in-context learning to improve dialogue safety
Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tur · 2023
Cited alongside, same era.
Play guessing game with LLM: indirect jailbreak attack with implicit clues
Zhiyuan Chang, Mingyang Li, Yi Liu, Junjie Wang, Qing Wang, and Yang Liu · 2024
Closest in time.
Leveraging the context through multi-round interactions for jailbreaking attacks
Yixin Cheng, Markos Georgopoulos, Volkan Cevher, and Grigorios G. Chrysos · 2024
Closest in time.
Combating adversarial attacks with multi-agent debate
Steffi Chern, Zhen Fan, and Andy Liu · 2024
Closest in time.
Breaking down the defenses: A comparative survey of attacks on large language models
Arijit Ghosh Chowdhury, Md Mofijul Islam, Vaibhav Kumar, Faysal Hossain Shezan, Vaibhav Kumar, Vinija Jain, and Aman Chadha · 2024
Closest in time.
Security and privacy challenges of large language models: A survey
Badhan Chandra Das, M. Hadi Amini, and Yanzhao Wu · 2024
Closest in time.
Pandora: Jailbreak gpts by retrieval augmented generation poisoning
Gelei Deng, Yi Liu, Kailong Wang, Yuekang Li, Tianwei Zhang, and Yang Liu · 2024
Closest in time.
The llama 3 herd of models, 2024
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, et al · 2024
Closest in time.
Red-teaming for generative AI: silver bullet or security theater?
Michael Feffer, Anusha Sinha, Zachary C. Lipton, and Hoda Heidari · 2024
Closest in time.
Improving adversarial transferability of visual-language pre-training models through collaborative multimodal interaction, 2024
Jiyuan Fu, Zhaoyu Chen, Kaixun Jiang, Haijing Guo, Jiafeng Wang, Shuyong Gao, and Wenqiang Zhang · 2024
Closest in time.
Coercing llms to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein · 2024
Closest in time.
Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation
Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, and Yu Zhang · 2024
Closest in time.
Agent smith: A single image can jailbreak one million multimodal LLM agents exponentially fast
Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Ye Wang, Jing Jiang, and Min Lin · 2024
Closest in time.
Cold-attack: Jailbreaking llms with stealthiness and controllability
Xingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin, and Bin Hu · 2024
Closest in time.
Towards safe and aligned large language models for medicine
Tessa Han, Aounon Kumar, Chirag Agarwal, and Himabindu Lakkaraju · 2024
Closest in time.
Jailbreaking proprietary large language models using word substitution cipher
Divij Handa, Advait Chirmule, Bimal G. Gajera, and Chitta Baral · 2024
Closest in time.
Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho · 2024
Closest in time.
Trustagent: Towards safe and trustworthy llm-based agents through agent constitution
Wenyue Hua, Xianjun Yang, Zelong Li, Wei Cheng, and Yongfeng Zhang · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam S. Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul F. Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
Defending large language models against jailbreak attacks via semantic smoothing
Jiabao Ji, Bairu Hou, Alexander Robey, George J. Pappas, Hamed Hassani, Yang Zhang, Eric Wong, and Shiyu Chang · 2024
Closest in time.
Haibo Jin, Ruoxi Chen, Andy Zhou, Jinyin Chen, Yang Zhang, and Haohan Wang · 2024
Closest in time.
Nevermind: Instruction override and moderation in large language models
Edward Kim · 2024
Closest in time.
Break the breakout: Reinventing LM defense against jailbreak attacks with self-refinement
Heegyu Kim, Sehyun Yuk, and Hyunsouk Cho · 2024
Closest in time.
Can llms recognize toxicity? structured toxicity investigation framework and semantic-based metric
Hyukhun Koh, Dohyung Kim, Minwoo Lee, and Kyomin Jung · 2024
Closest in time.
The ethics of interaction: Mitigating security threats in llms
Ashutosh Kumar, Sagarika Singh, Shiv Vignesh Murty, and Swathy Ragupathy · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea · 2024
Closest in time.
Using hallucinations to bypass gpt4’s filter
Benjamin Lemkin · 2024
Closest in time.
Mitigating the alignment tax of rlhf, 2024
Yong Lin, Hangyu Lin, Wei Xiong, Shizhe Diao, Jianmeng Liu, Jipeng Zhang, Rui Pan, Haoxiang Wang, Wenbin Hu, Hanning Zhang, Hanze Dong, Renjie Pi, Han Zhao, Nan Jiang, Heng Ji, Yuan Yao, and Tong Zhang · 2024
Closest in time.
A safe harbor for AI evaluation and red teaming
Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng Xin Yong, Suhas Kotha, Yi Zeng, Weiyan Shi, Xianjun Yang, Reid Southen, Alexander Robey, Patrick Chao, Diyi Yang, Ruoxi Jia, Daniel Kang, Sandy Pentland, Arvind Narayanan, Percy Liang, and Peter Henderson · 2024
Closest in time.
Ensuring safe and high-quality outputs: A guideline library approach for language models, 2024
Yi Luo, Zhenghao Lin, Yuhao Zhang, Jiashuo Sun, Chen Lin, Chengjin Xu, Xiangdong Su, Yelong Shen, Jian Guo, and Yeyun Gong · 2024
Closest in time.
PRP: propagating universal perturbations to attack large language model guard-rails
Neal Mangaokar, Ashish Hooda, Jihye Choi, Shreyas Chandrashekaran, Kassem Fawaz, Somesh Jha, and Atul Prakash · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David A. Forsyth, and Dan Hendrycks · 2024
Closest in time.
Adversarial text purification: A large language model approach for defense
Raha Moraffah, Shubh Khandelwal, Amrita Bhattacharjee, and Huan Liu · 2024
Closest in time.
Openai usage policies
OpenAI · 2024
Closest in time.
How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions
Lorenzo Pacchiardi, Alex James Chan, Sören Mindermann, Ilan Moscovitz, Alexa Yue Pan, Yarin Gal, Owain Evans, and Jan M. Brauner · 2024
Closest in time.
Attacking LLM watermarks by exploiting their strengths
Qi Pang, Shengyuan Hu, Wenting Zheng, and Virginia Smith · 2024
Closest in time.
Mapping llm security landscapes: A comprehensive stakeholder risk assessment proposal, 2024
Rahul Pankajakshan, Sumitra Biswal, Yuvaraj Govindarajulu, and Gilad Gressel · 2024
Closest in time.
Neural exec: Learning (and learning from) execution triggers for prompt injection attacks
Dario Pasquini, Martin Strohmeier, and Carmela Troncoso · 2024
Closest in time.
Mllm-protector: Ensuring mllm’s safety without hurting performance
Renjie Pi, Tianyang Han, Yueqi Xie, Rui Pan, Qing Lian, Hanze Dong, Jipeng Zhang, and Tong Zhang · 2024
Closest in time.
Visual adversarial examples jailbreak aligned large language models
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal · 2024
Closest in time.
Learning to poison large language models during instruction tuning
Yao Qiang, Xiangyu Zhou, Saleh Zare Zade, Mohammad Amin Roshani, Douglas Zytko, and Dongxiao Zhu · 2024
Closest in time.
Vision-llms can fool themselves with self-generated typographic attacks
Maan Qraitem, Nazia Tasnim, Piotr Teterwak, Kate Saenko, and Bryan A. Plummer · 2024
Closest in time.
Guardian: A multi-tiered defense architecture for thwarting prompt injection attacks on llms, 2024
Parijat Rai, Saumil Sood, Vijay K Madisetti, and Arshdeep Bahga · 2024
Closest in time.
Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot LLM assessment
Vyas Raina, Adian Liusie, and Mark J. F. Gales · 2024
Closest in time.
Towards red teaming in multimodal and multilingual translation
Christophe Ropers, David Dale, Prangthip Hansanti, Gabriel Mejia Gonzalez, Ivan Evtimov, Corinne Wong, Christophe Touret, Kristina Pereyra, Seohyun Sonia Kim, Cristian Canton-Ferrer, Pierre Andrews, and Marta R. Costa-jussà · 2024
Closest in time.
Immunization against harmful fine-tuning attacks
Domenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze, Jan Batzner, Hassan Sajjad, and Frank Rudzicz · 2024
Closest in time.
An early categorization of prompt injection attacks on large language models
Sippo Rossi, Alisia Marianne Michel, Raghava Rao Mukkamala, and Jason Bennett Thatcher · 2024
Closest in time.
Fast adversarial attacks on language models in one GPU minute
Vinu Sankar Sadasivan, Shoumik Saha, Gaurang Sriramanan, Priyatham Kattakinda, Atoosa Malemir Chegini, and Soheil Feizi · 2024
Closest in time.
The butterfly effect of altering prompts: How small changes and jailbreaks affect large language model performance, 2024
Abel Salinas and Fred Morstatter · 2024
Closest in time.
On the conversational persuasiveness of large language models: A randomized controlled trial, 2024
Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West · 2024
Closest in time.
Prompt stealing attacks against large language models
Zeyang Sha and Yang Zhang · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Closest in time.
SPML: A DSL for defending language models against prompt attacks
Reshabh K. Sharma, Vinayak Gupta, and Dan Grossman · 2024
Closest in time.
Attackeval: How to evaluate the effectiveness of jailbreak attacking on large language models
Dong Shu, Mingyu Jin, Suiyuan Zhu, Beichen Wang, Zihao Zhou, Chong Zhang, and Yongfeng Zhang · 2024
Closest in time.
PAL: proxy-guided black-box attack on large language models
Chawin Sitawarin, Norman Mu, David A. Wagner, and Alexandre Araujo · 2024
Closest in time.
A strongreject for empty jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer · 2024
Closest in time.
Trustllm: Trustworthiness in large language models, 2024
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, Joaquin Vanschoren, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen, Tianming Liu, Tianyi Zhou, William Wang, Xiang Li, Xiangliang Zhang, Xiao Wang, Xing Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, Yong Chen, and Yue Zhao · 2024
Closest in time.
Scaling behavior of machine translation with large language models under prompt injection attacks
Zhifan Sun and Antonio Valerio Miceli Barone · 2024
Closest in time.
Xuchen Suo · 2024
Closest in time.
Prioritizing safeguarding over autonomy: Risks of LLM agents for science
Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan, Yichi Zhang, Wangchunshu Zhou, Meng Qu, Yilun Zhao, Jian Tang, Zhuosheng Zhang, Arman Cohan, Zhiyong Lu, and Mark Gerstein · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, Andrea Tacchetti, Colin Gaffney, Samira Daruki, Olcan Sercinoglu, Zach Gleicher, Juliette Love, Paul Voigtlaender, et al · 2024
Closest in time.
Zephyr: Direct distillation of LM alignment
Lewis Tunstall, Edward Emanuel Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro Von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M Rush, and Thomas Wolf · 2024
Closest in time.
Neeraj Varshney, Pavel Dolin, Agastya Seth, and Chitta Baral · 2024
Closest in time.
Veil: A payload generation framework
Veil-Framework · 2024
Closest in time.
Uncovering mesa-optimization algorithms in transformers, 2024
Johannes von Oswald, Maximilian Schlegel, Alexander Meulemans, Seijin Kobayashi, Eyvind Niklasson, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Blaise Agüera y Arcas, Max Vladymyrov, Razvan Pascanu, and João Sacramento · 2024
Closest in time.
Assessing the brittleness of safety alignment via pruning and low-rank modifications, 2024
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson · 2024
Closest in time.
Gradient-based language model red teaming
Nevan Wichers, Carson Denison, and Ahmad Beirami · 2024
Closest in time.
Tradeoffs between alignment and helpfulness in language models
Yotam Wolf, Noam Wies, Dorin Shteyman, Binyamin Rothberg, Yoav Levine, and Amnon Shashua · 2024
Closest in time.
Universal prompt optimizer for safe text-to-image generation
Zongyu Wu, Hongcheng Gao, Yueze Wang, Xiang Zhang, and Suhang Wang · 2024
Closest in time.
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li · 2024
Closest in time.
Tastle: Distract large language models for automatic jailbreak attack
Zeguan Xiao, Yan Yang, Guanhua Chen, and Yun Chen · 2024
Closest in time.
Linkprompt: Natural and universal adversarial attacks on prompt-based language models, 2024
Yue Xu and Wenjie Wang · 2024
Closest in time.
Llm lies: Hallucinations are not bugs, but features as adversarial examples, 2024
Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu, Mu-Nan Ning, Yu-Yang Liu, and Li Yuan · 2024
Closest in time.
Toolsword: Unveiling safety issues of large language models in tool learning across three stages
Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, and Xuanjing Huang · 2024
Closest in time.
Daniel Wankit Yip, Aysan Esmradi, and Chun-Fai Chan · 2024
Closest in time.
R-judge: Benchmarking safety risk awareness for LLM agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, and Gongshen Liu · 2024
Closest in time.
Round trip translation defence against large language model jailbreaking attacks
Canaan Yung, Hadi Mohaghegh Dolatabadi, Sarah M. Erfani, and Christopher Leckie · 2024
Closest in time.
On prompt-driven safeguarding for large language models, 2024
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy M. Hospedales · 2024
Closest in time.