Fetching the paper…
Reading the bibliography…
Current research in adversarial robustness of LLMs focuses on discrete input manipulations in the natural language space, which can be directly transferred to closed-source models.
Teoria statistica delle classi e calcolo delle probabilita
Carlo Bonferroni · 1936
Earlier work this paper cites.
Individual comparisons by ranking methods
Frank Wilcoxon · 1992
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Universal adversarial perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard · 2017
Earlier work this paper cites.
Towards Deep Learning Models Resistant to Adversarial Attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao · 2019
Earlier work this paper cites.
Freelb: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
Deep learning for time series classification: a review
Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhassane Idoumghar, and Pierre-Alain Muller · 2019
Earlier work this paper cites.
Adversarial training for large neural language models
Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Detoxify
Laura Hanu and Unitary team · 2020
Earlier work this paper cites.
Time matters: Time-aware LSTMs for predictive business process monitoring
An Nguyen, Srijeet Chatterjee, Sven Weinzierl, Leo Schwinn, Martin Matzner, and Bjoern M. Eskofier · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Token-aware virtual adversarial training in natural language understanding
Linyang Li and Xipeng Qiu · 2021
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
CLIP: Cheap lipschitz training of neural networks
Leon Bungert, René Raab, Tim Roith, Leo Schwinn, and Daniel Tenbrinck · 2021
Earlier work this paper cites.
System design for a data-driven and explainable customer sentiment monitor using iot and enterprise data
An Nguyen, Stefan Foerstel, Thomas Kittler, Andrey Kurzyukov, Leo Schwinn, Dario Zanca, Tobias Hipp, Sun Da Jun, Michael Schrapp, Eva Rothgang, et al · 2021
Earlier work this paper cites.
Improved text classification via contrastive adversarial training
Lin Pan, Chung-Wei Hang, Avirup Sil, and Saloni Potdar · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Cited alongside, same era.
Improving robustness against real-world and worst-case distribution shifts through decision region quantification
Leo Schwinn, René Raab, An Nguyen, Dario Zanca, and Bjoern M. Eskofier · 2022
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Rain: Your language models can align themselves without finetuning
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Later among the works it cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Later among the works it cites.
Raising the bar for certified adversarial robustness with diffusion models
Thomas Altstidl, David Dobre, Björn Eskofier, Gauthier Gidel, and Leo Schwinn · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang · 2023
Cited alongside, same era.
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Cited alongside, same era.
Chatgpt, 2023
OpenAI · 2023
Cited alongside, same era.
Adversarial attacks and defenses in large language models: Old and new threats
Leo Schwinn, David Dobre, Stephan Günnemann, and Gauthier Gidel · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
Identifying and mitigating the security risks of generative ai
Clark Barrett, Brad Boyd, Ellie Burzstein, Nicholas Carlini, Brad Chen, Jihye Choi, Amrita Roy Chowdhury, Mihai Christodorescu, Anupam Datta, Soheil Feizi, et al · 2023
Cited alongside, same era.
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Later among the works it cites.
Catastrophic jailbreak of open-source LLMs via exploiting generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter · 2024
Closest in time.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2024
Closest in time.
Attacking large language models with projected gradient descent
Simon Geisler, Tom Wollschläger, MHI Abdalla, Johannes Gasteiger, and Stephan Günnemann · 2024
Closest in time.
An adversarial perspective on machine unlearning for ai safety
Jakub Łucki, Boyi Wei, Yangsibo Huang, Peter Henderson, Florian Tramèr, and Javier Rando · 2024
Closest in time.
Evaluating deep unlearning in large language models
Ruihan Wu, Chhavi Yadav, Russ Salakhutdinov, and Kamalika Chaudhuri · 2024
Closest in time.
Sos! soft prompt attack against open-source large language models
Ziqing Yang, Michael Backes, Yang Zhang, and Ahmed Salem · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Improving alignment and robustness with short circuiting
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al · 2024
Closest in time.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Closest in time.
Llama 3 model card
AI@Meta · 2024
Closest in time.
A strongreject for empty jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al · 2024
Closest in time.
Defending against unforeseen failure modes with latent adversarial training
Stephen Casper, Lennart Schulze, Oam Patel, and Dylan Hadfield-Menell · 2024
Closest in time.
A probabilistic perspective on unlearning and alignment for large language models
Yan Scholten, Stephan Günnemann, and Leo Schwinn · 2024
Closest in time.