Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) rely on vision encoders and Large Language Models (LLMs) to exhibit remarkable capabilities on various multi-modal tasks in the joint space of vision and language.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on. pp. 248–255. IEEE (2009)
2009
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13. pp. 740–755. Springer (2014)
2014
Earlier work this paper cites.
Malinowski, M., Fritz, M.: A multi-world approach to question answering about real-world scenes based on uncertain input. Advances in neural information processing systems 27
2014
Earlier work this paper cites.
Zhu, Y., Groth, O., Bernstein, M., Fei-Fei, L.: Visual7w: Grounded question answering in images. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4995–5004 (2016)
2016
Earlier work this paper cites.
Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad-cam: Visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE international conference on computer vision. pp. 618–626 (2017)
2017
Earlier work this paper cites.
Acharya, M., Kafle, K., Kanan, C.: Tallyqa: Answering complex counting questions. In: Proceedings of the AAAI conference on artificial intelligence. vol. 33, pp. 8076–8084 (2019)
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
Yang, L., Han, Y., Chen, X., Song, S., Dai, J., Huang, G.: Resolution adaptive networks for efficient inference. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 2369–2378 (2020)
2020
Earlier work this paper cites.
Conde, M.V., Turgutlu, K.: Clip-art: Contrastive pre-training for fine-grained art classification. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3956–3960 (2021)
2021
Earlier work this paper cites.
Goh, G., †, N.C., †, C.V., Carter, S., Petrov, M., Schubert, L., Radford, A., Olah, C.: Multimodal neurons in artificial neural networks. Distill (2021). https://doi.org/10.23915/distill.00030, https://distill.pub/2021/multimodal-neurons
2021
Earlier work this paper cites.
Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-language representation learning with noisy text supervision. In: International conference on machine learning. pp. 4904–4916. PMLR (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: Meila, M., Zhang, T. (eds.) Proceedings of the 38th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 139, pp. 8748–8763. PMLR (18–24 Jul 2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Yang, L., Jiang, H., Cai, R., Wang, Y., Song, S., Huang, G., Tian, Q.: Condensenet v2: Sparse feature reactivation for deep networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3569–3578 (2021)
2021
Earlier work this paper cites.
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al.: Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems 35
2022
Earlier work this paper cites.
Avrahami, O., Lischinski, D., Fried, O.: Blended diffusion for text-driven editing of natural images. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18208–18218 (2022)
2022
Earlier work this paper cites.
Cheng, H., Xu, K., Li, Z., Zhao, P., Wang, C., Lin, X., Kailkhura, B., Goldhahn, R.: More or less (mol): Defending against multiple perturbation attacks on deep neural networks through model ensemble and compression. In: 2022 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW). pp. 645–655. IEEE (2022)
2022
Earlier work this paper cites.
Gu, J., Tresp, V., Qin, Y.: Are vision transformers robust to patch perturbations? In: European Conference on Computer Vision. pp. 404–421. Springer (2022)
2022
Earlier work this paper cites.
Li, J., Li, D., Xiong, C., Hoi, S.: Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning. pp. 12888–12900. PMLR (2022)
2022
Earlier work this paper cites.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.: Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35
2022
Earlier work this paper cites.
Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., Lu, J.: Denseclip: Language-guided dense prediction with context-aware prompting. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18082–18091 (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al.: Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
Schwenk, D., Khandelwal, A., Clark, C., Marino, K., Mottaghi, R.: A-okvqa: A benchmark for visual question answering using world knowledge. In: European Conference on Computer Vision. pp. 146–162. Springer (2022)
2023
Later among the works it cites.
2023
Later among the works it cites.
Lu, D., Wang, Z., Wang, T., Guan, W., Gao, H., Zheng, F.: Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 102–111 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Zhang, J., Yi, Q., Sang, J.: Towards adversarial attack on vision-language pre-training models. In: Proceedings of the 30th ACM International Conference on Multimedia. pp. 5005–5013 (2022)
2022
Cited alongside, same era.
Zhang, R., Zhang, W., Fang, R., Gao, P., Li, K., Dai, J., Qiao, Y., Li, H.: Tip-adapter: Training-free adaption of clip for few-shot classification. In: European Conference on Computer Vision. pp. 493–510. Springer (2022)
2022
Cited alongside, same era.
Zhou, C., Loy, C.C., Dai, B.: Extract free dense labels from clip. In: European Conference on Computer Vision. pp. 696–712. Springer (2022)
2022
Cited alongside, same era.
Azuma, H., Matsui, Y.: Defense-prefix for preventing typographic attacks on clip. ICCV Workshop on Adversarial Robustness In the Real World (2023)
2023
Cited alongside, same era.
Dai, W., Li, J., Li, D., Tiong, A.M.H., Zhao, J., Wang, W., Li, B., Fung, P., Hoi, S.: Instructblip: Towards general-purpose vision-language models with instruction tuning (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Duan, J., Fan, Q., Cheng, H., Shi, X., Xu, K.: Improve video representation with temporal adversarial augmentation. Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI) (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Yang, L., Zheng, Z., Wang, J., Song, S., Huang, G., Li, F.: Adadet: An adaptive object detection system based on early-exit neural networks. IEEE Transactions on Cognitive and Developmental Systems 16
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
Cheng, H., Cao, J., Xiao, E., Sun, M., Xu, R.: Gaining the sparse rewards by exploring binary lottery tickets in spiking neural network. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (2024)
2024
Closest in time.
Cheng, H., Duan, J., Li, H., Zhang, L., Cao, J., Wang, P., Zhang, J., Xu, K., Xu, R.: Rbformer: Improve adversarial robustness of transformer by robust bias. British Machine Vision Conference (BMVC) (2024)
2024
Closest in time.
Duan, J., Cheng, H., Wang, S., Wang, C., Zavalny, A., Xu, R., Kailkhura, B., Xu, K.: Shifting attention to relevance: Towards the uncertainty estimation of large language models. The 62nd Annual Meeting of the Association for Computational Linguistics (ACL) (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Kong, F., Duan, J., Sun, L., Cheng, H., Xu, R., Shen, H., Zhu, X., Shi, X., Xu, K.: Act-diffusion: Efficient adversarial consistency training for one-step diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8890–8899 (2024)
2024
Closest in time.
2024
Closest in time.