Fetching the paper…
Reading the bibliography…
With the advent and widespread deployment of Multimodal Large Language Models (MLLMs), ensuring their safety has become increasingly critical.
Adversarial Machine Learning at Scale
Kurakin, A., Goodfellow, I. J., and Bengio, S · 2017
Earlier work this paper cites.
Adversarial Attacks Are Reversible With Natural Supervision
Mao, C., Chiquier, M., Wang, H., Yang, J., and Vondrick, C · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Diffusion models for adversarial purification
Nie, W., Guo, B., Huang, Y., Xiao, C., Vahdat, A., and Anandkumar, A · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2022
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries, 2023
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic
Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R · 2023
Earlier work this paper cites.
Chen, Y., Sikka, K., Cogswell, M., Ji, H., and Divakaran, A · 2023
Earlier work this paper cites.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., Li, B., Luo, P., Lu, T., Qiao, Y., and Dai, J · 2023
Earlier work this paper cites.
Opencompass: A universal evaluation platform for foundation models
Contributors, O · 2023
Earlier work this paper cites.
How Robust is Google’s Bard to Adversarial Image Attacks?
Dong, Y., Chen, H., Chen, J., Fang, Z., Yang, X., Zhang, Y., Tian, Y., Su, H., and Zhu, J · 2023
Earlier work this paper cites.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., Wu, Y., and Ji, R · 2023
Earlier work this paper cites.
Chain of Thought Prompt Tuning in Vision Language Models
Ge, J., Luo, H., Qian, S., Gan, Y., Fu, J., and Zhan, S · 2023
Earlier work this paper cites.
Figstep: Jailbreaking large vision-language models via typographic visual prompts, 2023
Gong, Y., Ran, D., Liu, J., Wang, C., Cong, T., Wang, A., Duan, S., and Wang, X · 2023
Earlier work this paper cites.
Han, D., Jia, X., Bai, Y., Gu, J., Liu, Y., and Cao, X · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M · 2023
Earlier work this paper cites.
Large Language Models as Automated Aligners for benchmarking Vision-Language Models
Ji, Y., Ge, C., Kong, W., Xie, E., Liu, Z., Li, Z., and Luo, P · 2023
Earlier work this paper cites.
Mistral 7b, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Earlier work this paper cites.
Silkie: Preference Distillation for Large Visual Language Models
Li, L., Xie, Z., Li, M., Chen, S., Wang, P., Chen, L., Yang, Y., Wang, B., and Kong, L · 2023
Earlier work this paper cites.
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Lin, B., Zhu, B., Ye, Y., Ning, M., Jin, P., and Yuan, L · 2023
Earlier work this paper cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2023
Earlier work this paper cites.
GPT-4v(ision) as a social media analysis engine
Lyu, H., Huang, J., Zhang, D., Yu, Y., Mou, X., Pan, J., Yang, Z., Wei, Z., and Luo, J · 2023
Earlier work this paper cites.
Gpt-4 technical report, 2023
openai team · 2023
Earlier work this paper cites.
Visual Adversarial Examples Jailbreak Aligned Large Language Models
Qi, X., Huang, K., Panda, A., Henderson, P., Wang, M., and Mittal, P · 2023
Earlier work this paper cites.
On the adversarial robustness of multi-modal foundation models
Schlarmann, C., and Hein, M · 2023
Cited alongside, same era.
Role-play with large language models, 2023
Shanahan, M., McDonell, K., and Reynolds, L · 2023
Cited alongside, same era.
Survey of vulnerabilities in large language models revealed by adversarial attacks
Shayegani, E., Mamun, M. A. A., Fu, Y., Zaree, P., Dong, Y., and Abu-Ghazaleh, N · 2023
Cited alongside, same era.
"do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models, 2023
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Cited alongside, same era.
Reflexion: language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. R., and Yao, S · 2023
Cited alongside, same era.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.
A Survey on Hallucination in Large Vision-Language Models
Liu, H., Xue, W., Chen, Y., Chen, D., Zhao, X., Wang, K., Hou, L., Li, R., and Peng, W · 2024
Closest in time.
Democratizing fine-grained visual recognition with large language models
Liu, M., Roy, S., Li, W., Zhong, Z., Sebe, N., and Ricci, E · 2024
Closest in time.
Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing
Liu, S., Nie, W., Wang, C., Lu, J., Qiao, Z., Liu, L., Tang, J., Xiao, C., and Anandkumar, A · 2024
Closest in time.
AgentBench: Evaluating LLMs as Agents
Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., and Tang, J · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sun, Z., Shen, S., Cao, S., Liu, H., Li, C., Shen, Y., Gan, C., Gui, L.-Y., Wang, Y.-X., Yang, Y., Keutzer, K., and Darrell, T · 2023
Cited alongside, same era.
Multi-party chat: Conversational agents in group settings with humans and models, 2023
Wei, J., Shuster, K., Szlam, A., Weston, J., Urbanek, J., and Komeili, M · 2023
Cited alongside, same era.
Skywork: A More Open Bilingual Foundation Model
Wei, T., Zhao, L., Zhang, L., Zhu, B., Wang, L., Yang, H., Li, B., Cheng, C., Lü, W., Hu, R., Li, C., Yang, L., Luo, X., Wu, X., Liu, L., Cheng, W., Cheng, P., Zhang, J., Zhang, X., Lin, L., Wang, X., Ma, Y., Dong, C., Sun, Y., Chen, Y., Peng, Y., Liang, X., Yan, S., Fang, H., and Zhou, Y · 2023
Cited alongside, same era.
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J · 2023
Cited alongside, same era.
A Survey on Multimodal Large Language Models
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E · 2023
Cited alongside, same era.
Woodpecker: Hallucination Correction for Multimodal Large Language Models
Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., and Chen, E · 2023
Cited alongside, same era.
Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., Sun, M., et al · 2023
Cited alongside, same era.
Closest in time.
Mm-safetybench: A benchmark for safety evaluation of multimodal large language models, 2024
Liu, X., Zhu, Y., Gu, J., Lan, Y., Yang, C., and Qiao, Y · 2024
Closest in time.
Llm discussion: Enhancing the creativity of large language models via discussion framework and role-play, 2024
Lu, L.-C., Chen, S.-J., Pai, T.-M., Yu, C.-H., yi Lee, H., and Sun, S.-H · 2024
Closest in time.
Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks, 2024
Luo, W., Ma, S., Liu, X., Guo, X., and Xiao, C · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D · 2024
Closest in time.
Llama 2 - acceptable use policy
Meta AI · 2024
Closest in time.
A Comprehensive Overview of Large Language Models
Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., and Mian, A · 2024
Closest in time.
Jailbreaking Attack against Multimodal Large Language Model
Niu, Z., Ren, H., Gao, X., Hua, G., and Jin, R · 2024
Closest in time.
Usage policies - openai
OpenAI · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T. P., Alayrac, J., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., Antonoglou, I., Anil, R., Borgeaud, S., Dai, A. M., Millican, K., Dyer, E., Glaese, M., Sottiaux, T., Lee, B., Viola, F., Reynolds, M., Xu, Y., Molloy, J., Chen, J., Isard, M., Barham, P., Hennigan, T., McIlroy, R., Johnson, M., Schalkwyk, J., Collins, E., Rutherford, E., Moreira, E., Ayoub, K., Goel, M., Meyer, C., Thornton, G., Yang, Z., Michalewski, H., Abbas, Z., Schucher, N., Anand, A., Ives, R., Keeling, J., Lenc, K., Haykal, S., Shakeri, S., Shyam, P., Chowdhery, A., Ring, R., Spencer, S., Sezener, E., and et al · 2024
Closest in time.
Zero shot VLMs for hate meme detection: Are we there yet?
Rizwan, N., Bhaskar, P., Das, M., Majhi, S. S., Saha, P., and Mukherjee, A · 2024
Closest in time.
Lamp: When large language models meet personalization, 2024
Salemi, A., Mysore, S., Bendersky, M., and Zamani, H · 2024
Closest in time.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E., Dong, Y., and Abu-Ghazaleh, N · 2024
Closest in time.
Rolecraft-glm: Advancing personalized role-playing in large language models, 2024
Tao, M., Liang, X., Shi, T., Yu, L., and Xie, Y · 2024
Closest in time.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., Truong, S. T., Arora, S., Mazeika, M., Hendrycks, D., Lin, Z., Cheng, Y., Koyejo, S., Song, D., and Li, B · 2024
Closest in time.
Wang, Y., Liu, X., Li, Y., Chen, M., and Xiao, C · 2024
Closest in time.
Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models, 2024
Wang, Z. M., Peng, Z., Que, H., Liu, J., Zhou, W., Wu, Y., Guo, H., Gan, R., Ni, Z., Yang, J., Zhang, M., Zhang, Z., Ouyang, W., Xu, K., Huang, S. W., Fu, J., and Peng, J · 2024
Closest in time.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking, 2024
Xu, N., Wang, F., Zhou, B., Li, B. Z., Xiao, C., and Chen, M · 2024
Closest in time.
Large language models as optimizers, 2024
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2024
Closest in time.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms, 2024
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.
MM-LLMs: Recent Advances in MultiModal Large Language Models
Zhang, D., Yu, Y., Li, C., Dong, J., Su, D., Chu, C., and Yu, D · 2024
Closest in time.
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Timothy, H · 2024
Closest in time.