Fetching the paper…
Reading the bibliography…
Despite promising performance on open-source large vision-language models (LVLMs), transfer-based targeted attacks often fail against closed-source commercial LVLMs.
Towards evaluating the robustness of neural networks
N. Carlini and D. Wagner · 2017
Earlier work this paper cites.
Nips 2017: Defense against adversarial attack
A. K, B. Hamner, and I. Goodfellow · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
D. P. Kingma and J. Ba · 2017
Earlier work this paper cites.
Delving into transferable adversarial examples and black-box attacks
Y. Liu, X. Chen, C. Liu, and D. Song · 2017
Earlier work this paper cites.
Boosting adversarial attacks with momentum
Y. Dong, F. Liao, T. Pang, H. Su, J. Zhu, X. Hu, and J. Li · 2018
Earlier work this paper cites.
Black-box adversarial attacks with limited queries and information
A. Ilyas, L. Engstrom, A. Athalye, and J. Lin · 2018
Earlier work this paper cites.
Adversarial examples in the physical world
A. Kurakin, I. J. Goodfellow, and S. Bengio · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2018
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Earlier work this paper cites.
Query-efficient black-box adversarial attacks guided by a transfer-based prior
Y. Dong, S. Cheng, T. Pang, H. Su, and J. Zhu · 2021
Earlier work this paper cites.
Openclip, July 2021
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Visualgpt: Data-efficient adaptation of pretrained language models for image captioning
J. Chen, H. Guo, K. Yi, B. Li, and M. Elhoseiny · 2022
Earlier work this paper cites.
Scaling up vision-language pre-training for image captioning
X. Hu, Z. Gan, J. Wang, Z. Yang, Z. Liu, Y. Lu, and L. Wang · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
How robust is google’s bard to adversarial image attacks?
Y. Dong, H. Chen, J. Chen, Z. Fang, X. Yang, Y. Zhang, Y. Tian, H. Su, and J. Zhu · 2023
Cited alongside, same era.
Gptscore: Evaluate as you desire
J. Fu, S.-K. Ng, Z. Jiang, and P. Liu · 2023
The revolution of multimodal large language models: A survey
D. Caffagni, F. Cocchi, L. Barsellotti, N. Moratelli, S. Sarto, L. Baraldi, M. Cornia, and R. Cucchiara · 2024
Later among the works it cites.
Rethinking model ensemble in transfer-based adversarial attacks
H. Chen, Y. Zhang, Y. Dong, X. Yang, H. Su, and J. Zhu · 2024
Later among the works it cites.
Efficient generation of targeted and transferable adversarial examples for vision-language models via diffusion models
Q. Guo, S. Pang, X. Jia, Y. Liu, and Q. Guo · 2024
Later among the works it cites.
Enhancing advanced visual reasoning ability of large language models
Z. Li, D. Liu, C. Zhang, H. Wang, T. Xue, and W. Cai · 2024
Later among the works it cites.
A survey of multimodel large language models
Z. Liang, Y. Xu, Y. Hong, P. Shang, Q. Wang, Q. Fu, and K. Liu · 2024
Later among the works it cites.
Questioning, answering, and captioning for zero-shot detailed image caption
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Cited alongside, same era.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
Crepe: Can vision-language foundation models reason compositionally?
Z. Ma, J. Hong, M. O. Gul, M. Gandhi, I. Gao, and R. Krishna · 2023
Cited alongside, same era.
Image captioning for effective use of language models in knowledge-based visual question answering
A. Salaberria, G. Azkune, O. L. de Lacalle, A. Soroa, and E. Agirre · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al · 2023
Cited alongside, same era.
Image captioners are scalable vision learners too
M. Tschannen, M. Kumar, A. Steiner, X. Zhai, N. Houlsby, and L. Beyer · 2023
Cited alongside, same era.
On evaluating adversarial robustness of large vision-language models
Y. Zhao, T. Pang, C. Du, X. Yang, C. Li, N.-M. M. Cheung, and M. Lin · 2023
Cited alongside, same era.
D.-T. Luu, V.-T. Le, and D. M. Vo · 2024
Later among the works it cites.
Enhancing visual question answering through question-driven image captions as prompts
Ö. Özdemir and E. Akagündüz · 2024
Later among the works it cites.
Cross-modal retrieval: a systematic review of methods and future directions
T. Wang, F. Li, L. Zhu, J. Li, Z. Zhang, and H. T. Shen · 2024
Later among the works it cites.
Mm-llms: Recent advances in multimodal large language models
D. Zhang, Y. Yu, J. Dong, C. Li, D. Su, C. Chu, and D. Yu · 2024
Later among the works it cites.
J. Zhang, J. Ye, X. Ma, Y. Li, Y. Yang, J. Sang, and D.-Y. Yeung · 2024
Later among the works it cites.
Introducing claude 3.5 sonnet, 2024
Anthropic · 2025
Closest in time.
Generalizing from simple to hard visual reasoning: Can we mitigate modality imbalance in vlms?
S. Park, A. Panigrahi, Y. Cheng, D. Yu, A. Goyal, and S. Arora · 2025
Closest in time.
Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks
J. Wu, M. Zhong, S. Xing, Z. Lai, Z. Liu, Z. Chen, W. Wang, X. Zhu, L. Lu, T. Lu, et al · 2025
Closest in time.