Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) have shown significant progress in responding well to visual-instructions from users.
A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra, “Visual dialog,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 326–335
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
G. Zheng, S. Mukherjee, X. L. Dong, and F. Li, “Opentag: Open attribute value extraction from product profiles,” in Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining , 2018, pp. 1049–1058
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Gao, J. Lanchantin, M. L. Soffa, and Y. Qi, “Black-box generation of adversarial text sequences to evade deep learning classifiers,” in 2018 IEEE Security and Privacy Workshops (SPW) , 2018, pp. 50–56
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Ren, Y. Deng, K. He, and W. Che, “Generating natural language adversarial examples through probability weighted word saliency,” in Proceedings of the 57th annual meeting of the association for computational linguistics , 2019, pp. 1085–1097
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
D. Jin, Z. Jin, J. T. Zhou, and P. Szolovits, “Is bert really robust? a strong baseline for natural language attack on text classification and entailment,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, 2020, pp. 8018–8025
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning , 2021, pp. 8748–8763
2021
Earlier work this paper cites.
T. Maho, T. Furon, and E. Le Merrer, “Surfree: a fast surrogate-free black-box attack,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 10 430–10 439
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual question answering using world knowledge,” in European Conference on Computer Vision , 2022, pp. 146–162
2022
Earlier work this paper cites.
Y. Shi, Y. Han, Y.-a. Tan, and X. Kuang, “Decision-based black-box attack against vision transformers via patch-wise adversarial removal,” Advances in Neural Information Processing Systems , vol. 35, pp. 12 921–12 933, 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez et al. , “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” See https://vicuna. lmsys. org (accessed 14 April 2023) , vol. 2, no. 3, p. 6, 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Advances in Neural Information Processing Systems , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Zhang, S. Lai, Y. Wang, Z. Da, Y. Dun, and X. Qian, “Scgnet: Shifting and cascaded group network,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 33, no. 9, pp. 4997–5008, 2023
2023
Cited alongside, same era.
H. Kuang, H. Liu, Y. Wu, and R. Ji, “Semantically consistent visual representation for adversarial robustness,” IEEE Transactions on Information Forensics and Security , vol. 18, pp. 5608–5622, 2023
2023
Cited alongside, same era.
T. Bai, J. Zhao, and B. Wen, “Guided adversarial contrastive distillation for robust students,” IEEE Transactions on Information Forensics and Security , pp. 1–1, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
C. Schlarmann and M. Hein, “On the adversarial robustness of multi-modal foundation models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3677–3685
2023
Cited alongside, same era.
2023
Later among the works it cites.
N. Bitton-Guetta, Y. Bitton, J. Hessel, L. Schmidt, Y. Elovici, G. Stanovsky, and R. Schwartz, “Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional images,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2616–2627
2023
Later among the works it cites.
H. Zhang, L. Xu, S. Lai, W. Shao, N. Zheng, P. Luo, Y. Qiao, and K. Zhang, “Open-vocabulary animal keypoint detection with semantic-feature matching,” International Journal of Computer Vision , vol. 132, no. 12, pp. 5741–5758, 2024
2024
Closest in time.
2024
Closest in time.
H. Zhang, Y. Ma, K. Zhang, N. Zheng, and S. Lai, “Fmgnet: An efficient feature-multiplex group network for real-time vision task,” Pattern Recognition , p. 110698, 2024
2024
Closest in time.
H. Zhang, Y. Dun, Y. Pei, S. Lai, C. Liu, K. Zhang, and X. Qian, “Hf-hrnet: A simple hardware friendly high-resolution network,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 34, no. 8, pp. 7699–7711, 2024
2024
Closest in time.
H. Kuang, H. Liu, X. Lin, and R. Ji, “Defense against adversarial attacks using topology aligning adversarial training,” IEEE Transactions on Information Forensics and Security , vol. 19, pp. 3659–3673, 2024
2024
Closest in time.
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. A. Forsyth, and D. Hendrycks, “Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,” in Forty-first International Conference on Machine Learning , 2024
2024
Closest in time.
2024
Closest in time.
Y. Zhao, T. Pang, C. Du, X. Yang, C. Li, N.-M. M. Cheung, and M. Lin, “On evaluating adversarial robustness of large vision-language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 296–26 306
2024
Closest in time.
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
B. Li, Y. Ge, Y. Ge, G. Wang, R. Wang, R. Zhang, and Y. Shan, “Seed-bench: Benchmarking multimodal large language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 299–13 308
2024
Closest in time.
S. Chen, J. Gu, Z. Han, Y. Ma, P. Torr, and V. Tresp, “Benchmarking robustness of adaptation methods on pre-trained vision-language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
L. Li, H. Guan, J. Qiu, and M. Spratling, “One prompt word is enough to boost adversarial robustness for pre-trained vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 24 408–24 419
2024
Closest in time.
C. Schlarmann, N. D. Singh, F. Croce, and M. Hein, “Robust clip: Unsupervised adversarial fine-tuning of vision embeddings for robust large vision-language models,” in Forty-first International Conference on Machine Learning , 2024
2024
Closest in time.
L. Bailey, E. Ong, S. Russell, and S. Emmons, “Image hijacks: Adversarial images can control generative models at runtime,” in Forty-first International Conference on Machine Learning , 2024
2024
Closest in time.
X. Cui, A. Aparcedo, Y. K. Jang, and S.-N. Lim, “On the robustness of large multimodal models against image adversarial attacks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 24 625–24 634
2024
Closest in time.
2024
Closest in time.
N. Carlini, M. Nasr, C. A. Choquette-Choo, M. Jagielski, I. Gao, P. W. W. Koh, D. Ippolito, F. Tramer, and L. Schmidt, “Are aligned neural networks adversarially aligned?” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.