Fetching the paper…
Reading the bibliography…
Attracted by the impressive power of Multimodal Large Language Models (MLLMs), the public is increasingly utilizing them to improve the efficiency of daily work.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Protest activity detection and perceived violence estimation from social media images
Donghyeon Won, Zachary C. Steinert-Threlkeld, et al · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Multi-modal sarcasm detection in Twitter with hierarchical fusion model
Yitao Cai, Huiyu Cai, and Xiaojun Wan · 2019
Earlier work this paper cites.
Kvqa: Knowledge-aware visual question answering
Sanket Shah, Anand Mishra, et al · 2019
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine · 2020
Earlier work this paper cites.
Multimodal meme dataset (MultiOFF) for identifying offensive content in image and text
Shardul Suryawanshi, Bharathi Raja Chakravarthi, et al · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, et al · 2021
Earlier work this paper cites.
Detecting harmful memes and their targets
Shraman Pramanick, Dimitar Dimitrov, Rituparna Mukherjee, Shivam Sharma, Md. Shad Akhtar, Preslav Nakov, and Tanmoy Chakraborty · 2021
Earlier work this paper cites.
MOMENTA: A multimodal framework for detecting harmful memes and their targets
Shraman Pramanick, Shivam Sharma, Dimitar Dimitrov, Md. Shad Akhtar, Preslav Nakov, and Tanmoy Chakraborty · 2021
Earlier work this paper cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, et al · 2021
Earlier work this paper cites.
SemEval-2022 task 5: Multimedia automatic misogyny identification
Elisabetta Fersini, Francesca Gasparini, Giulia Rizzi, Aurora Saibene, Berta Chulvi, Paolo Rosso, Alyssa Lees, and Jeffrey Sorensen · 2022
Earlier work this paper cites.
Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs
Eugene Bagdasaryan, Tsung-Yin Hsieh, Ben Nassi, and Vitaly Shmatikov · 2023
Earlier work this paper cites.
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
Luke Bailey, Euan Ong, Stuart Russell, and Scott Emmons · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt · 2023
Earlier work this paper cites.
Can pre-trained vision and language models answer visual information-seeking questions?
Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, and Ming-Wei Chang · 2023
Cited alongside, same era.
Can language models be instructed to protect personal information?
Yang Chen, Ethan Mendes, Sauvik Das, Wei Xu, and Alan Ritter · 2023
Cited alongside, same era.
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran · 2023
Cited alongside, same era.
How robust is google’s bard to adversarial image attacks?
Yinpeng Dong, Huanran Chen, Jiawei Chen, Zhengwei Fang, Xiao Yang, Yichi Zhang, Yu Tian, Hang Su, and Jun Zhu · 2023
Cited alongside, same era.
Misusing Tools in Large Language Models With Visual Adversarial Examples
How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs
Haoqin Tu, Chenhang Cui, et al · 2023
Later among the works it cites.
Detecting and correcting hate speech in multimodal memes with large visual language model
Minh-Hao Van and Xintao Wu · 2023
Later among the works it cites.
Adventures of trustworthy vision-language models: A survey
Mayank Vatsa, Anubhooti Jain, and Richa Singh · 2023
Later among the works it cites.
Tovilag: Your visual-language generative model is also an evildoer
Xinpeng Wang, Xiaoyuan Yi, et al · 2023
Later among the works it cites.
Jailbreaking gpt-4v via self-adversarial attacks with system prompts
Yuanwei Wu, Xiang Li, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaohan Fu, Zihan Wang, Shuheng Li, Rajesh K. Gupta, Niloofar Mireshghallah, Taylor Berg-Kirkpatrick, and Earlence Fernandes · 2023
Cited alongside, same era.
Figstep: Jailbreaking large vision-language models via typographic visual prompts
Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Cited alongside, same era.
Large Language Models as Automated Aligners for benchmarking Vision-Language Models
Yuanfeng Ji, Chongjian Ge, Weikai Kong, Enze Xie, Zhengying Liu, Zhengguo Li, and Ping Luo · 2023
Cited alongside, same era.
Query-relevant images jailbreak large multi-modal models
Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang, and Yu Qiao · 2023
Cited alongside, same era.
Visual adversarial examples jailbreak aligned large language models
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Mengdi Wang, and Prateek Mittal · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Cited alongside, same era.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Rose Kirk, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Bridge the gap between CV and NLP! a gradient-based textual adversarial attack framework
Lifan Yuan, YiChi Zhang, et al · 2023
Later among the works it cites.
Safetybench: Evaluating the safety of large language models with multiple choice questions
Zhexin Zhang, Leqi Lei, et al · 2023
Later among the works it cites.
Red teaming visual language models
Mukai Li, Lei Li, Yuwei Yin, Masood Ahmed, Zhenguang Liu, and Qi Liu · 2024
Closest in time.
Goat-bench: Safety insights to large multimodal models through meme-based social abuse
Hongzhan Lin, Ziyang Luo, Bo Wang, Ruichao Yang, and Jing Ma · 2024
Closest in time.
An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models
Haochen Luo, Jindong Gu, Fengyuan Liu, and Philip Torr · 2024
Closest in time.
Mllm-protector: Ensuring mllm’s safety without hurting performance
Renjie Pi, Tianyang Han, Yueqi Xie, Rui Pan, Qing Lian, Hanze Dong, Jipeng Zhang, and Tong Zhang · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, et al · 2024
Closest in time.
Inferaligner: Inference-time alignment for harmlessness through cross-model guidance
Pengyu Wang, Dong Zhang, et al · 2024
Closest in time.
Llava- ϕ \phi : Efficient multi-modal assistant with small language model
Yichen Zhu, Minjie Zhu, et al · 2024
Closest in time.