Fetching the paper…
Reading the bibliography…
Large Visual Language Model\textbfs (VLMs) such as GPT-4V have achieved remarkable success in generating comprehensive and nuanced responses.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Zhang, W. E., Sheng, Q. Z., Alhazmi, A., and Li, C · 2020
Earlier work this paper cites.
Model extraction and adversarial transferability, your bert is vulnerable!
He, X., Lyu, L., Xu, Q., and Sun, L · 2021
Earlier work this paper cites.
Using adversarial attacks to reveal the statistical bias in machine reading comprehension models
Lin, J., Zou, J., and Ding, N · 2021
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2021
Earlier work this paper cites.
A survey on universal adversarial attack
Zhang, C., Benz, P., Lin, C., Karjauv, A., Wu, J., and Kweon, I. S · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., Long, D. X., Tan, J. Q., Joty, S., and Hoque, E · 2022
Earlier work this paper cites.
Infographicvqa
Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., and Jawahar, C · 2022
Earlier work this paper cites.
Introducing chatgpt, 2022
OpenAI · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., et al · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Perspective api, 2023
Google · 2023
Cited alongside, same era.
Catastrophic jailbreak of open-source llms via exploiting generation
Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D · 2023
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models
Lapid, R., Langberg, R., and Sipper, M · 2023
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., and Xing, X · 2023
Later among the works it cites.
Investigating copyright issues of diffusion models under practical scenarios
Zhang, Y., Tzun, T. T., Hern, L. W., Wang, H., and Kawaguchi, K · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Many-shot jailbreaking
Anil, C., Durmus, E., Sharma, M., Benton, J., Kundu, S., Batson, J., Rimsky, N., Tong, M., Mu, J., Ford, D., et al · 2024
Closest in time.
Can llms’ tuning methods work in medical multimodal domain?
Chen, J., Jiang, Y., Yang, D., Li, M., Wei, J., Qian, Z., and Zhang, L · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-step jailbreaking privacy attacks on chatgpt
Li, H., Guo, D., Fan, W., Xu, M., and Song, Y · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2023
Cited alongside, same era.
Fairness-guided few-shot prompting for large language models
Ma, H., Zhang, C., Bian, Y., Liu, L., Zhang, Z., Zhao, P., Zhang, S., Fu, H., Hu, Q., and Wu, B · 2023
Cited alongside, same era.
Visual language integration: A survey and open challenges
Park, S.-M. and Kim, Y.-G · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Cited alongside, same era.
Flowchartqa: the first large-scale benchmark for reasoning over flowcharts
Tannert, S., Feighelstein, M. G., Bogojeska, J., Shtok, J., Arbelle, A., Staar, P. W., Schumann, A., Kuhn, J., and Karlinsky, L · 2023
Cited alongside, same era.
Closest in time.
Egothink: Evaluating first-person perspective thinking capability of vision-language models
Cheng, S., Guo, Z., Wu, J., Fang, K., Li, P., Liu, H., and Liu, Y · 2024
Closest in time.
Large language models as optimizers
Chengrun, Y., Xuezhi, W., Yifeng, L., Hanxiao, L., V, L. Q., Denny, Z., and Xinyun, C · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Hu, S., Tu, Y., Han, X., He, C., Cui, G., Long, X., Zheng, Z., Fang, Y., Huang, Y., Zhao, W., et al · 2024
Closest in time.
Automatic jailbreaking of the text-to-image generative ai systems
Kim, M., Lee, H., Gong, B., Zhang, H., and Hwang, S. J · 2024
Closest in time.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., and Fried, D · 2024
Closest in time.
Visual adversarial examples jailbreak aligned large language models
Qi, X., Huang, K., Panda, A., Henderson, P., Wang, M., and Mittal, P · 2024
Closest in time.
Bridge the modality and capacity gaps in vision-language model selection
Yi, C., Zhan, D.-C., and Ye, H.-J · 2024
Closest in time.
Discovering universal semantic triggers for text-to-image synthesis
Zhai, S., Wang, W., Li, J., Dong, Y., Su, H., and Shen, Q · 2024
Closest in time.
On evaluating adversarial robustness of large vision-language models
Zhao, Y., Pang, T., Du, C., Yang, X., Li, C., Cheung, N.-M. M., and Lin, M · 2024
Closest in time.
Is the system message really important to jailbreaks in large language models?
Zou, X., Chen, Y., and Li, K · 2024
Closest in time.