Fetching the paper…
Reading the bibliography…
Vision-Language Models like GPT-4, LLaVA, and CogVLM have surged in popularity recently due to their impressive performance in several vision-language tasks.
Vovk V., Gammerman A., Shafer G. Algorithmic learning in a random world. – New York : Springer, 2005. – T. 29
2005
Earlier work this paper cites.
2009
Earlier work this paper cites.
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: ´ Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740– 755. Springer, 2014
2014
Earlier work this paper cites.
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision, pp. 2641– 2649, 2015
2015
Earlier work this paper cites.
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., & Farhadi, A. (2016). A diagram is worth a dozen images. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14 (pp. 235-251). Springer International Publishing
2016
Earlier work this paper cites.
Kazemzadeh, Sahar, et al. ”ReferItGame: Referring to Objects in Photographs of Natural Scenes.” EMNLP 2014. Yu, Licheng, et al. ”Modeling Context in Referring Expressions.” ECCV 2016
2016
Earlier work this paper cites.
Guo, C., Pleiss, G., Sun, Y., & Weinberger K. Q. (2017). On Calibration of Modern Neural Networks. In Proceedings of the 34th International Conference on Machine Learning, PMLR 70:1321-1330
2017
Earlier work this paper cites.
Goyal, Yash, et al. ”Making the v in vqa matter: Elevating the role of image understanding in visual question answering.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017
2017
Earlier work this paper cites.
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P. Nocaps: Novel object captioning at scale. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 8948–8957, 2019
2019
Earlier work this paper cites.
Hudson, Drew A., and Christopher D. Manning. ”Gqa: A new dataset for real-world visual reasoning and compositional question answering.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019
2019
Earlier work this paper cites.
Singh, Amanpreet, et al. ”Towards vqa models that can read.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019
2019
Earlier work this paper cites.
Sadinle, M., Lei, J., & Wasserman, L. (2019). Least ambiguous set-valued classifiers with bounded error levels. Journal of the American Statistical Association, 114(525), 223-234
2019
Earlier work this paper cites.
Romano Y., Sesia M., Candes E. Classification with valid and adaptive coverage //Advances in Neural Information Processing Systems. – 2020. – T. 33. – C. 3581-3591
2020
Earlier work this paper cites.
Maltoudoglou, L., Paisios, A., & Papadopoulos, H. (2020, August). BERT-based conformal predictor for sentiment analysis. In Conformal and Probabilistic Prediction and Applications (pp. 269-284). PMLR
2020
Earlier work this paper cites.
Romano, Y., Sesia, M., & Candes, E. (2020). Classification with valid and adaptive coverage. Advances in Neural Information Processing Systems, 33, 3581-3591
2020
Earlier work this paper cites.
Giovannotti, P., & Gammerman, A. (2021, September). Transformer-based conformal predictors for paraphrase detection. In Conformal and Probabilistic Prediction and Applications (pp. 243-265). PMLR
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Zhao, B., Yu, S., Ma, W., Yu, M., Mei, S., Wang, A., He, J., Yuille, A. & Kortylewski, A. OOD-CV: A Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images. (2022)
2022
Earlier work this paper cites.
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K., Zhu, S., Tafjord, O., Clark, P. & Kalyan, A. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. (2022)
2022
Earlier work this paper cites.
Angelopoulos, A. & Bates, S. A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. (2022)
2022
Earlier work this paper cites.
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X. & Cao, Y. EVA: Exploring the Limits of Masked Visual Representation Learning at Scale. (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Dey, N., Ding, J., Ferrell, J., Kapper, C., Lovig, M., Planchon, E., & Williams, J. P. (2022). Conformal Prediction for Text Infilling and Part-of-Speech Prediction. The New England Journal of Statistics in Data Science, 1(1), 69-83
2022
Cited alongside, same era.
Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., Chen, K. & Lin, D. MMBench: Is Your Multi-modal Model an All-around Player?. (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Li, B., Wang, R., Wang, G., Ge, Y., Ge, Y. & Shan, Y. SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension. (2023)
2023
Cited alongside, same era.
Zhang, D., Li, S., Zhang, X., Zhan, J., Wang, P., Zhou, Y. & Qiu, X. SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities. The 2023 Conference On Empirical Methods In Natural Language Processing
2023
Later among the works it cites.
2023
Later among the works it cites.
Jian, Y., Gao, C. & Vosoughi, S. Bootstrapping Vision-Language Learning with Decoupled Language Pre-training. (2023)
2023
Later among the works it cites.
Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A. & Taşırlar, S. Introducing our Multimodal Models. (2023), https://www.adept.ai/blog/fuyu-8b
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J., Yu, J., Soricut, R., Schalkwyk, J., Dai, A., Hauth, A., Millican, K., Silver, D., Petrov, S., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T., Lazaridou, A., Firat, O., Molloy, J., Isard, M., Barham, P., Hennigan, T., Lee, B., Viola, F., Reynolds, M., Xu, Y., Doherty, R., Collins, E., Meyer, C., Rutherford, E., Moreira, E., Ayoub, K., Goel, M., Tucker, G., Piqueras, E., Krikun, M., Barr, I., Savinov, N., Danihelka, I., Roelofs, B., White, A., Andreassen, A., Glehn, T., Yagati, L., Kazemi, M., Gonzalez, L., Khalman, M., Sygnowski, J., Frechette, A., Smith, C., Culp, L., Proleev, L., Luan, Y., Chen, X., Lottes, J., Schucher, N., Lebron, F., Rrustemi, A., Clay, N., Crone, et al. Gemini: A Family of Highly Capable Multimodal Models. (2023)
2023
Cited alongside, same era.
Awadalla, A., Gao, I., Gardner, J., Hessel, J., Hanafy, Y., Zhu, W., Marathe, K., Bitton, Y., Gadre, S., Sagawa, S., Jitsev, J., Kornblith, S., Koh, P., Ilharco, G., Wortsman, M. & Schmidt, L. OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models. (2023)
2023
Cited alongside, same era.
Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
2023
Cited alongside, same era.
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C. & Zhou, J. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Li, J., Li, D., Savarese, S. & Hoi, S. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. (2023)
2023
Later among the works it cites.
Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mohammed, O., Patra, B., Liu, Q., Aggarwal, K., Chi, Z., Bjorck, J., Chaudhary, V., Som, S., Song, X. & Wei, F. Language Is Not All You Need: Aligning Perception with Language Models. (2023)
2023
Later among the works it cites.
Li, B., Zhang, P., Yang, J., Zhang, Y., Pu, F. & Liu, Z. OtterHD: A High-Resolution Multi-modality Model. (2023)
2023
Later among the works it cites.
Team, I. Internlm: A multilingual language model with progressively enhanced capabilities. 2023-01-06)[2023-09-27]. Https://github. Com/InternLM/InternLM
2023
Later among the works it cites.
Fontana M., Zeni G., Vantini S. Conformal prediction: a unified review of theory and new challenges //Bernoulli. – 2023. – T. 29. – 1. – C. 1-23
2023
Later among the works it cites.
2023
Later among the works it cites.
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Li, C., Xu, Y., Chen, H., Tian, J., Qi, Q., Zhang, J. & Huang, F. mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality. (2023)
2023
Later among the works it cites.
Zhang, P., Dong, X., Wang, B., Cao, Y., Xu, C., Ouyang, L., Zhao, Z., Duan, H., Zhang, S., Ding, S., Zhang, W., Yan, H., Zhang, X., Li, W., Li, J., Chen, K., He, C., Zhang, X., Qiao, Y., Lin, D. & Wang, J. InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition. (2023)
2023
Later among the works it cites.
Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y. & Bai, X. Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models. (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
Zhang, D., Yu, Y., Li, C., Dong, J., Su, D., Chu, C. & Yu, D. MM-LLMs: Recent Advances in MultiModal Large Language Models. (2024)
2024
Closest in time.
Lin, B., Tang, Z., Ye, Y., Cui, J., Zhu, B., Jin, P., Huang, J., Zhang, J., Ning, M. & Yuan, L. MoE-LLaVA: Mixture of Experts for Large Vision-Language Models. (2024)
2024
Closest in time.