Fetching the paper…
Reading the bibliography…
Large vision-language models (LVLMs) have demonstrated their incredible capability in image understanding and response generation.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: CVPR. pp. 248–255. IEEE (2009)
2009
Earlier work this paper cites.
Biggio, B., Corona, I., Maiorca, D., Nelson, B., Šrndić, N., Laskov, P., Giacinto, G., Roli, F.: Evasion attacks against machine learning at test time. In: Machine Learning and Knowledge Discovery in Databases. pp. 387–402. Springer Berlin Heidelberg, Berlin, Heidelberg (2013)
2013
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: ECCV. pp. 740–755. Springer (2014)
2014
Earlier work this paper cites.
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., Fergus, R.: Intriguing properties of neural networks. In: ICLR. pp. 1–10. http://OpenReview.net, Banff, AB, Canada (2014)
2014
Earlier work this paper cites.
Goodfellow, I.J., Shlens, J., Szegedy, C.: Explaining and harnessing adversarial examples. In: ICLR. pp. 1–10. http://OpenReview.net, San Diego, CA, USA (2015)
2015
Earlier work this paper cites.
Carlini, N., Wagner, D.: Towards evaluating the robustness of neural networks. In: IEEE Symposium on Security and Privacy. pp. 39–57. IEEE, San Jose, CA, USA (2017)
2017
Earlier work this paper cites.
Chen, P.Y., Zhang, H., Sharma, Y., Yi, J., Hsieh, C.J.: Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In: ACM Workshop on Artificial Intelligence and Security. pp. 15–26. Association for Computing Machinery, New York, NY, USA (2017)
2017
Earlier work this paper cites.
Kurakin, A., Goodfellow, I., Bengio, S.: Adversarial machine learning at scale. In: ICLR. pp. 1–17. http://OpenReview.net, Toulon, France (2017)
2017
Earlier work this paper cites.
Liu, Y., Chen, X., Liu, C., Song, D.: Delving into transferable adversarial examples and black-box attacks. In: ICLR (2017), https://openreview.net/forum?id=Sys6GJqxl
2017
Earlier work this paper cites.
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., Vladu, A.: Towards deep learning models resistant to adversarial attacks. In: ICLR. pp. 1–28. http://OpenReview.net, Vancouver, BC, Canada (2017)
2017
Earlier work this paper cites.
Papernot, N., McDaniel, P., Goodfellow, I., Jha, S., Celik, Z.B., Swami, A.: Practical black-box attacks against machine learning. In: ACM on Asia Conference on Computer and Communications Security. pp. 506–519. Association for Computing Machinery, New York, NY, USA (2017)
2017
Earlier work this paper cites.
Akhtar, N., Mian, A.: Threat of adversarial attacks on deep learning in computer vision: A survey. IEEE Access 6
2018
Earlier work this paper cites.
Chen, H., Zhang, H., Chen, P.Y., Yi, J., Hsieh, C.J.: Attacking visual language grounding with adversarial examples: A case study on neural image captioning. In: ACL (2018)
2018
Earlier work this paper cites.
Dong, Y., Liao, F., Pang, T., Su, H., Zhu, J., Hu, X., Li, J.: Boosting adversarial attacks with momentum. In: CVPR. pp. 9185–9193. IEEE, Salt Lake City, UT, USA (2018)
2018
Earlier work this paper cites.
Ilyas, A., Engstrom, L., Athalye, A., Lin, J.: Black-box adversarial attacks with limited queries and information. In: ICML. pp. 2137–2146. PMLR, Stockholmsmässan, Stockholm SWEDEN (2018)
2018
Earlier work this paper cites.
Xu, X., Chen, X., Liu, C., Rohrbach, A., Darrell, T., Song, D.: Fooling vision and language models despite localization and attention mechanism. In: CVPR. pp. 4951–4961 (2018)
2018
Earlier work this paper cites.
Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: CVPR (2018)
2018
Earlier work this paper cites.
Li, C., Gao, S., Deng, C., Xie, D., Liu, W.: Cross-modal learning with adversarial samples. In: NeurIPS (2019)
2019
Earlier work this paper cites.
Gao, L., Zhang, Q., Song, J., Liu, X., Shen, H.T.: Patch-wise attack for fooling deep neural network. In: ECCV. pp. 307–322. Springer (2020)
2020
Earlier work this paper cites.
Liu, A., Huang, T., Liu, X., Xu, Y., Ma, Y., Chen, X., Maybank, S.J., Tao, D.: Spatiotemporal attacks for embodied agents. In: ECCV. pp. 122–138. Springer (2020)
2020
Cited alongside, same era.
Tang, R., Ma, C., Zhang, W.E., Wu, Q., Yang, X.: Semantic equivalent adversarial data augmentation for visual question answering. In: ECCV. pp. 437–453. Springer (2020)
2020
Cited alongside, same era.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: ICLR (2021)
2021
Cited alongside, same era.
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., Hoi, S.C.H.: Align before fuse: Vision and language representation learning with momentum distillation. In: NeurIPS. pp. 9694–9705 (2021)
2021
Cited alongside, same era.
2023
Closest in time.
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., Cao, Y.: Eva: Exploring the limits of masked visual representation learning at scale. In: CVPR. pp. 19358–19369 (2023)
2023
Closest in time.
2023
Closest in time.
Google: Google bard (2023), https://bard.google.com/
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: ICML. pp. 8748–8763. PMLR (2021)
2021
Cited alongside, same era.
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al.: Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
Chen, J., Guo, H., Yi, K., Li, B., Elhoseiny, M.: Visualgpt: Data-efficient adaptation of pretrained language models for image captioning. In: CVPR. pp. 18030–18040 (June 2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
OpenAI: ChatGPT. https://openai.com/blog/chatgpt (2022)
2022
Cited alongside, same era.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR. pp. 10684–10695 (2022)
2022
Cited alongside, same era.
Zhai, X., Kolesnikov, A., Houlsby, N., Beyer, L.: Scaling vision transformers. In: CVPR. pp. 12104–12113 (2022)
2022
Cited alongside, same era.
Zhang, J., Yi, Q., Sang, J.: Towards adversarial attack on vision-language pre-training models. In: ACM MM. pp. 5005–5013 (2022)
2022
Cited alongside, same era.
2023
Closest in time.
Li, J., Li, D., Savarese, S., Hoi, S.: Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In: ICML (2023)
2023
Closest in time.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. In: NeurIPS (2023)
2023
Closest in time.
Lu, D., Wang, Z., Wang, T., Guan, W., Gao, H., Zheng, F.: Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models. In: ICCV. pp. 102–111 (2023)
2023
Closest in time.
OpenAI: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
2023
Closest in time.
OpenAI: Gpt-4v(ision) system card (2023), https://openai.com/research/gpt-4v-system-card
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Yin, Z., Ye, M., Zhang, T., Du, T., Zhu, J., Liu, H., Chen, J., Wang, T., Ma, F.: Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models. In: NeurIPS (2023)
2023
Closest in time.
Yu, L., Rieser, V.: Adversarial textual robustness on visual dialog. In: ACL. pp. 3422–3438 (2023)
2023
Closest in time.
Zhao, Y., Pang, T., Du, C., Yang, X., Li, C., Cheung, N.M., Lin, M.: On evaluating adversarial robustness of large vision-language models. In: NeurIPS (2023)
2023
Closest in time.
Zhou, Z., Hu, S., Li, M., Zhang, H., Zhang, Y., Jin, H.: Advclip: Downstream-agnostic adversarial examples in multimodal contrastive learning. In: ACM MM. pp. 6311–6320 (2023)
2023
Closest in time.
2023
Closest in time.
Shayegani, E., Dong, Y.: Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models. In: ICLR (2024)
2024
Closest in time.