Fetching the paper…
Reading the bibliography…
Compared with Large Language Models (LLMs), Large Vision-Language Models (LVLMs) can also accept images as input, thus showcasing more interesting emergent capabilities and demonstrating impressive performance on various vision-language tasks.
Bigham, J.P., Jayant, C., Ji, H., Little, G., Miller, A., Miller, R.C., Miller, R., Tatarowicz, A., White, B., White, S., Yeh, T.: Vizwiz: nearly real-time answers to visual questions. In: Proceedings of the 23nd Annual ACM Symposium on User Interface Software and Technology. p. 333–342 (2010)
2010
Earlier work this paper cites.
Singh, A., Natarjan, V., Shah, M., Jiang, Y., Chen, X., Parikh, D., Rohrbach, M.: Towards vqa models that can read. In: IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR). pp. 8317–8326 (2019)
2019
Earlier work this paper cites.
Gao, T., Fisch, A., Chen, D.: Making pre-trained language models better few-shot learners. In: Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (ICML) (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: Meila, M., Zhang, T. (eds.) International Conference on Machine Learning (ICML) (2021)
2021
Earlier work this paper cites.
Yao, Y., Zhang, A., Zhang, Z., Liu, Z., Chua, T., Sun, M.: CPT: colorful prompt tuning for pre-trained vision-language models. CoRR (2021)
2021
Earlier work this paper cites.
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al.: Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems 35
2022
Earlier work this paper cites.
Dong, B., Zhou, P., Yan, S., Zuo, W.: LPT: long-tailed prompt tuning for image classification. CoRR (2022)
2022
Earlier work this paper cites.
Du, Y., Wei, F., Zhang, Z., Shi, M., Gao, Y., Li, G.: Learning to prompt for open-vocabulary object detection with vision-language model. In: IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) (2022)
2022
Earlier work this paper cites.
Fahes, M., Vu, T., Bursuc, A., Pérez, P., de Charette, R.: Pøda: Prompt-driven zero-shot domain adaptation. CoRR (2022)
2022
Earlier work this paper cites.
Ganaie, M.A., Hu, M., Malik, A.K., Tanveer, M., Suganthan, P.N.: Ensemble deep learning: A review. Eng. Appl. Artif. Intell. 115
2022
Earlier work this paper cites.
Jia, M., Tang, L., Chen, B., Cardie, C., Belongie, S.J., Hariharan, B., Lim, S.: Visual prompt tuning. In: European Conference on Computer Vision (ECCV) (2022)
2022
Earlier work this paper cites.
Jia, M., Tang, L., Chen, B., Cardie, C., Belongie, S.J., Hariharan, B., Lim, S.: Visual prompt tuning. In: European Conference on Computer Vision (ECCV) (2022)
2022
Earlier work this paper cites.
Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., Iwasawa, Y.: Large language models are zero-shot reasoners. Advances in neural information processing systems 35
2022
Earlier work this paper cites.
Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., Iwasawa, Y.: Large language models are zero-shot reasoners. Advances in neural information processing systems 35
2022
Earlier work this paper cites.
Niu, H., Li, H., Zhao, F., Li, B.: Domain-unified prompt representations for source-free domain generalization. CoRR (2022)
2022
Earlier work this paper cites.
Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., Lu, J.: Denseclip: Language-guided dense prediction with context-aware prompting. In: IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) (2022)
2022
Earlier work this paper cites.
Shen, S., Yang, S., Zhang, T., Zhai, B., Gonzalez, J.E., Keutzer, K., Darrell, T.: Multitask vision-language prompt tuning. CoRR (2022)
2022
Earlier work this paper cites.
Shu, M., Nie, W., Huang, D., Yu, Z., Goldstein, T., Anandkumar, A., Xiao, C.: Test-time prompt tuning for zero-shot generalization in vision-language models. CoRR (2022)
2022
Earlier work this paper cites.
Shu, M., Nie, W., Huang, D., Yu, Z., Goldstein, T., Anandkumar, A., Xiao, C.: Test-time prompt tuning for zero-shot generalization in vision-language models. In: Conference on Neural Information Processing Systems 2022, NeurIPS (2022)
2022
Earlier work this paper cites.
Wang, W., Cao, Y., Zhang, J., Tao, D.: FP-DETR: detection transformer advanced by fully pre-training. In: International Conference on Learning Representations (ICLR) (2022)
2022
Earlier work this paper cites.
Wang, Y., Huang, Z., Hong, X.: S-prompts learning with pre-trained transformers: An occam’s razor for domain incremental learning. CoRR (2022)
2022
Earlier work this paper cites.
Wang, Z., Zhang, Z., Lee, C., Zhang, H., Sun, R., Ren, X., Su, G., Perot, V., Dy, J.G., Pfister, T.: Learning to prompt for continual learning. In: CVPR (2022)
2022
Earlier work this paper cites.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al.: Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35
2022
Earlier work this paper cites.
Wu, C.H., Motamed, S., Srivastava, S., la Torre, F.D.: Generative visual prompt: Unifying distributional control of pre-trained generative models. CoRR (2022)
2022
Earlier work this paper cites.
Zhang, Z., Zhou, Y., Zhao, X., Che, T., Lyu, L.: Prompt certified machine unlearning with randomized gradient smoothing and quantization. In: Conference on Neural Information Processing Systems (NeurlPS) (2022)
2022
Earlier work this paper cites.
Zhou, K., Yang, J., Loy, C.C., Liu, Z.: Conditional prompt learning for vision-language models. In: IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) (2022)
2022
Earlier work this paper cites.
Zhou, K., Yang, J., Loy, C.C., Liu, Z.: Conditional prompt learning for vision-language models. In: IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) (2022)
2022
Earlier work this paper cites.
Zhou, K., Yang, J., Loy, C.C., Liu, Z.: Learning to prompt for vision-language models. Int. J. Comput. Vis. (2022)
2022
Cited alongside, same era.
Zhou, K., Yang, J., Loy, C.C., Liu, Z.: Learning to prompt for vision-language models. Int. J. Comput. Vis. (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
Asai, A., Wu, Z., Wang, Y., Sil, A., Hajishirzi, H.: Self-rag: Learning to retrieve, generate, and critique through self-reflection. CoRR (2023)
2023
Later among the works it cites.
Pan, T., Tang, L., Wang, X., Shan, S.: Tokenize anything via prompting. CoRR (2023)
2023
Later among the works it cites.
Reddy, G.: The mechanistic basis of data dependence and abrupt learning in an in-context classification task. International Conference on Learning Representations (ICLR) (2023)
2023
Later among the works it cites.
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., Yao, S.: Reflexion: language agents with verbal reinforcement learning. In: Conference on Neural Information Processing Systems (NeurlPS) (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Burns, C., Izmailov, P., Kirchner, J.H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., Wu, J.: Weak-to-strong generalization: Eliciting strong capabilities with weak supervision (2023)
2023
Cited alongside, same era.
Chen, X., Lin, M., Schärli, N., Zhou, D.: Teaching large language models to self-debug. CoRR (2023)
2023
Cited alongside, same era.
Chowdhury, S., Nag, S., Manocha, D.: Apollo : Unified adapter and prompt learning for vision language models. In: Conference on Empirical Methods in Natural Language Processing, EMNLP (2023)
2023
Cited alongside, same era.
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., Hoi, S.: InstructBLIP: Towards general-purpose vision-language models with instruction tuning. In: Thirty-seventh Conference on Neural Information Processing Systems (2023), https://openreview.net/forum?id=vvoWPYqZJA
2023
Cited alongside, same era.
Darcet, T., Oquab, M., Mairal, J., Bojanowski, P.: Vision transformers need registers. CoRR (2023)
2023
Cited alongside, same era.
Du, Y., Li, S., Torralba, A., Tenenbaum, J.B., Mordatch, I.: Improving factuality and reasoning in language models through multiagent debate. CoRR (2023)
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Wang, T., Zhang, J., Fei, J., Zheng, H., Tang, Y., Li, Z., Gao, M., Zhao, S.: Caption anything: Interactive image description with diverse multimodal controls (2023)
2023
Later among the works it cites.
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., Xu, J., Xu, B., Li, J., Dong, Y., Ding, M., Tang, J.: Cogvlm: Visual expert for pretrained language models (2023)
2023
Later among the works it cites.
Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., Duan, N.: Visual chatgpt: Talking, drawing and editing with visual foundation models. CoRR (2023)
2023
Later among the works it cites.
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., Gao, J.: Set-of-mark prompting unleashes extraordinary visual grounding in GPT-4V. CoRR (2023)
2023
Later among the works it cites.
Yang, L., Wang, Y., Li, X., Wang, X., Yang, J.: Fine-grained visual prompting. In: Conference on Neural Information Processing Systems (NeurlPS) (2023)
2023
Later among the works it cites.
Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.C., Liu, Z., Wang, L.: The dawn of lmms: Preliminary explorations with gpt-4v(ision) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., Wang, L.: Mm-vet: Evaluating large multimodal models for integrated capabilities (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Zeng, A., Attarian, M., Ichter, B., Choromanski, K.M., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M.S., Sindhwani, V., Lee, J., Vanhoucke, V., Florence, P.: Socratic models: Composing zero-shot multimodal reasoning with language. In: International Conference on Learning Representations (ICLR) (2023)
2023
Later among the works it cites.
Zhang, A., Ji, W., Chua, T.: Next-chat: An LMM for chat, detection and segmentation. CoRR (2023)
2023
Later among the works it cites.
Zheng, C., Liu, Z., Xie, E., Li, Z., Li, Y.: Progressive-hint prompting improves reasoning in large language models. CoRR (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Gandelsman, Y., Efros, A.A., Steinhardt, J.: Interpreting CLIP’s image representation via text-based decomposition. In: International Conference on Learning Representations (ICLR) (2024)
2024
Closest in time.
Sahoo, P., Singh, A.K., Saha, S., Jain, V., Mondal, S., Chadha, A.: A systematic survey of prompt engineering in large language models: Techniques and applications. CoRR (2024)
2024
Closest in time.
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., Yao, S.: Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Wang, W., Ren, Y., Luo, H., Li, T., Yan, C., Chen, Z., Wang, W., Li, Q., Lu, L., Zhu, X., Qiao, Y., Dai, J.: The all-seeing project v2: Towards general relation comprehension of the open world (2024)
2024
Closest in time.
Zhang, Y., Ma, Z., Gao, X., Shakiah, S., Gao, Q., Chai, J.: Groundhog: Grounding large language models to holistic segmentation (2024)
2024
Closest in time.
Zhao, X., Yang, X., Pang, T., Du, C., Li, L., Wang, Y.X., Wang, W.Y.: Weak-to-strong jailbreaking on large language models (2024)
2024
Closest in time.