Fetching the paper…
Reading the bibliography…
The advent of large language models (LLMs) has revolutionized the field of natural language processing, yet they might be attacked to produce harmful content.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 1999
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Specializing word embeddings (for parsing) by information bottleneck
Xiang Lisa Li and Jason Eisner · 2019
Earlier work this paper cites.
Bottlesum: Unsupervised and self-supervised sentence summarization using the information bottleneck principle
Peter West, Ari Holtzman, Jan Buys, and Yejin Choi · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi · 2020
Earlier work this paper cites.
Improving question answering model robustness with synthetic adversarial data generation
Max Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel, Pontus Stenetorp, and Douwe Kiela · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Interpretable and generalizable graph learning via stochastic attention mechanism
Siqi Miao, Mia Liu, and Pan Li · 2022
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang · 2022
Earlier work this paper cites.
Recent advances in natural language processing via large pre-trained language models: A survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt · 2023
Cited alongside, same era.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2024
Closest in time.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2024
Closest in time.
Defending against alignment-breaking attacks via robustly aligned llm
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen · 2024
Closest in time.
Defending large language models against jailbreak attacks via semantic smoothing
Jiabao Ji, Bairu Hou, Alexander Robey, George J Pappas, Hamed Hassani, Yang Zhang, Eric Wong, and Shiyu Chang · 2024
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Cited alongside, same era.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Cited alongside, same era.
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Rusheb Shah, Soroush Pour, Arush Tagade, Stephen Casper, Javier Rando, et al · 2023
Cited alongside, same era.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Large language model unlearning
Yuanshun Yao, Xiaojun Xu, and Yang Liu · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Cited alongside, same era.
Beyond memorization: Violating privacy via inference with large language models
Robin Staab, Mark Vero, Mislav Balunović, and Martin Vechev · 2024
Closest in time.
Multilingual jailbreak challenges in large language models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2024
Closest in time.
Llm self defense: By self examination, llms know they are being tricked
Alec Helbling, Mansi Phute, Matthew Hull, and Duen Horng Chau · 2024
Closest in time.
Explaining time series via contrastive and locally sparse perturbations
Zichuan Liu, Yingying Zhang, Tianchun Wang, Zefan Wang, Dongsheng Luo, Mengnan Du, Min Wu, Yi Wang, Chunlin Chen, Lunting Fan, et al · 2024
Closest in time.
Timex++: Learning time-series explanations with information bottleneck
Zichuan Liu, Tianchun Wang, Jimeng Shi, Xu Zheng, Zhuomin Chen, Lei Song, Wenqian Dong, Jayantha Obeysekera, Farhad Shirani, and Dongsheng Luo · 2024
Closest in time.
Factorized explainer for graph neural networks
Rundong Huang, Farhad Shirani, and Dongsheng Luo · 2024
Closest in time.
Easyjailbreak: A unified framework for jailbreaking large language models
Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia, Yingshuang Gu, Mingxu Chai, Fukang Zhu, Caishuang Huang, Shihan Dou, Zhiheng Xi, et al · 2024
Closest in time.
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang · 2024
Closest in time.
Weak-to-strong jailbreaking on large language models
Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, and William Yang Wang · 2024
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.
Curiosity-driven red-teaming for large language models
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang, Aldo Pareja, James R. Glass, Akash Srivastava, and Pulkit Agrawal · 2024
Closest in time.