Fetching the paper…
Reading the bibliography…
Visual instruction tuning (VIT) has emerged as a crucial technique for enabling multi-modal large language models (MLLMs) to follow user instructions adeptly.
Marti, U.-V., Bunke, H.: The iam-database: an english sentence database for offline handwriting recognition. International Journal on Document Analysis and Recognition, 39–46 (2002) https://doi.org/10.1007/s100320200071
2002
Earlier work this paper cites.
Green, S., Heer, J., Manning, C.D.: The efficacy of human post-editing for language translation. In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, pp. 439–448 (2013)
2013
Earlier work this paper cites.
Medhat, W., Hassan, A., Korashy, H.: Sentiment analysis algorithms and applications: A survey. Ain Shams engineering journal 5
2014
Earlier work this paper cites.
Dong, D., Wu, H., He, W., Yu, D., Wang, H.: Multi-task learning for multiple language translation. In: Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 1723–1732 (2015)
2015
Earlier work this paper cites.
Ren, M., Kiros, R., Zemel, R.: Exploring models and data for image question answering. arXiv: Learning,arXiv: Learning (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., Farhadi, A.: A diagram is worth a dozen images. In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14, pp. 235–251 (2016). Springer
2016
Earlier work this paper cites.
Zhu, Y., Groth, O., Bernstein, M., Fei-Fei, L.: Visual7w: Grounded question answering in images. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016). https://doi.org/10.1109/cvpr.2016.540 . http://dx.doi.org/10.1109/cvpr.2016.540
2016
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Duan, N., Tang, D., Chen, P., Zhou, M.: Question generation for question answering. In: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 866–874 (2017)
2017
Earlier work this paper cites.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6904–6913 (2017)
2017
Earlier work this paper cites.
Kahou, S., Michalski, V., Atkinson, e.: Figureqa: An annotated figure dataset for visual reasoning. arXiv: Computer Vision and Pattern Recognition,arXiv: Computer Vision and Pattern Recognition (2017)
2017
Earlier work this paper cites.
Zhong, V., Xiong, C., Socher, R.: Seq2sql: Generating structured queries from natural language using reinforcement learning. Cornell University - arXiv,Cornell University - arXiv (2017)
2017
Earlier work this paper cites.
Iyyer, M., Yih, W.-t., Chang, M.-W.: Search-based neural structured learning for sequential question answering. In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2017). https://doi.org/10.18653/v1/p17-1167 . http://dx.doi.org/10.18653/v1/p17-1167
2017
Earlier work this paper cites.
Johnson, J., Hariharan, B., Maaten, L., Fei-Fei, L., Zitnick, C.L., Girshick, R.: Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017). https://doi.org/10.1109/cvpr.2017.215 . http://dx.doi.org/10.1109/cvpr.2017.215
2017
Earlier work this paper cites.
Kembhavi, A., Seo, M., Schwenk, D., Choi, J., Farhadi, A., Hajishirzi, H.: Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017). https://doi.org/10.1109/cvpr.2017.571 . http://dx.doi.org/10.1109/cvpr.2017.571
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Gurari, D., Li, Q., Stangl, A.J., Guo, A., Lin, C., Grauman, K., Luo, J., Bigham, J.P.: Vizwiz grand challenge: Answering visual questions from blind people. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3608–3617 (2018)
2018
Earlier work this paper cites.
Lau, J.J., Gayen, S., Ben Abacha, A., Demner-Fushman, D.: A dataset of clinically generated visual questions and answers about radiology images. Scientific Data (2018) https://doi.org/10.1038/sdata.2018.251
2018
Earlier work this paper cites.
Kafle, K., Price, B., Cohen, S., Kanan, C.: Dvqa: Understanding data visualizations via question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5648–5656 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al
2019
Earlier work this paper cites.
Marino, K., Rastegari, M., Farhadi, A., Mottaghi, R.: Ok-vqa: A visual question answering benchmark requiring external knowledge. In: Proceedings of the IEEE/cvf Conference on Computer Vision and Pattern Recognition, pp. 3195–3204 (2019)
2019
Earlier work this paper cites.
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., Rohrbach, M.: Towards vqa models that can read. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8317–8326 (2019)
2019
Earlier work this paper cites.
Hudson, D.A., Manning, C.D.: Gqa: A new dataset for real-world visual reasoning and compositional question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6700–6709 (2019)
2019
Earlier work this paper cites.
Goyal, Y., Khot, T., Agrawal, A., Summers-Stay, D., Batra, D., Parikh, D.: Making the v in vqa matter: Elevating the role of image understanding in visual question answering. International Journal of Computer Vision, 398–414 (2019) https://doi.org/10.1007/s11263-018-1116-0
2019
Earlier work this paper cites.
Acharya, M., Kafle, K., Kanan, C.: Tallyqa: Answering complex counting questions. Proceedings of the AAAI Conference on Artificial Intelligence, 8076–8084 (2019) https://doi.org/10.1609/aaai.v33i01.33018076
2019
Earlier work this paper cites.
Marino, K., Rastegari, M., Farhadi, A., Mottaghi, R.: Ok-vqa: A visual question answering benchmark requiring external knowledge. Cornell University - arXiv,Cornell University - arXiv (2019)
2019
Earlier work this paper cites.
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., Rohrbach, M.: Towards vqa models that can read. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019). https://doi.org/10.1109/cvpr.2019.00851 . http://dx.doi.org/10.1109/cvpr.2019.00851
2019
Earlier work this paper cites.
Biten, A.F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., Karatzas, D.: Scene text visual question answering. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4291–4301 (2019)
2019
Earlier work this paper cites.
Mishra, A., Shekhar, S., Singh, A.K., Chakraborty, A.: Ocr-vqa: Visual question answering by reading text in images. In: 2019 International Conference on Document Analysis and Recognition (ICDAR) (2019). https://doi.org/10.1109/icdar.2019.00156 . http://dx.doi.org/10.1109/icdar.2019.00156
2019
Earlier work this paper cites.
Methani, N., Ganguly, P., Khapra, M., Kumar, P.: Plotqa: Reasoning over scientific plots. Cornell University - arXiv,Cornell University - arXiv (2019)
2019
Earlier work this paper cites.
Zhang, C., Gao, F., Jia, B., Zhu, Y., Zhu, S.-C.: Raven: A dataset for relational and analogical visual reasoning. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019). https://doi.org/10.1109/cvpr.2019.00546 . http://dx.doi.org/10.1109/cvpr.2019.00546
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
2020
Earlier work this paper cites.
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., et al
2020
Earlier work this paper cites.
Kiela, D., Hamed, F., Mohan, A., Goswami, V., Singh, A., Ringshia, P., Testuggine, D.: The hateful memes challenge: Detecting hate speech in multimodal memes. Cornell University - arXiv,Cornell University - arXiv (2020)
2020
Earlier work this paper cites.
Pont-Tuset, J., Uijlings, J., Changpinyo, S., Soricut, R., Ferrari, V.: Connecting vision and language with localized narratives. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16, pp. 647–664 (2020). Springer
2020
Earlier work this paper cites.
Sidorov, O., Hu, R., Rohrbach, M., Singh, A.: Textcaps: a dataset for image captioning with reading comprehension. Cornell University - arXiv,Cornell University - arXiv (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Zhang, X., Sun, X., Luo, Y., Ji, J., Zhou, Y., Wu, Y., Huang, F., Ji, R.: Rstnet: Captioning with adaptive attention on visual and non-visual words. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15465–15474 (2021)
2021
Earlier work this paper cites.
Mathew, M., Karatzas, D., Jawahar, C.: Docvqa: A dataset for vqa on document images. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 2200–2209 (2021)
2021
Earlier work this paper cites.
Wang, B., Li, G., Zhou, X., Chen, Z., Grossman, T., Li, Y.: Screen2words: Automatic mobile ui summarization with multimodal learning. In: The 34th Annual ACM Symposium on User Interface Software and Technology, pp. 498–510 (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Chen, Z., Chen, W., Smiley, C., Shah, S., Borova, I., Langdon, D., Moussa, R., Beane, M., Huang, T.-Y., Routledge, B., Wang, W.: Finqa: A dataset of numerical reasoning over financial data. Cornell University - arXiv,Cornell University - arXiv (2021)
2021
Cited alongside, same era.
Dai, W., Li, J., Li, D., Tiong, A.M.H., Zhao, J., Wang, W., Li, B., Fung, P., Hoi, S.: InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al
2022
Cited alongside, same era.
Ma, Y., Xu, G., Sun, X., Yan, M., Zhang, J., Ji, R.: X-clip: End-to-end multi-grained contrastive learning for video-text retrieval. In: Proceedings of the 30th ACM International Conference on Multimedia, pp. 638–647 (2022)
2022
Cited alongside, same era.
Ma, Y., Ji, J., Sun, X., Zhou, Y., Wu, Y., Huang, F., Ji, R.: Knowing what it is: Semantic-enhanced dual attention transformer. IEEE Transactions on Multimedia 25
2022
Cited alongside, same era.
Li, J., Li, D., Xiong, C., Hoi, S.: Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning, pp. 12888–12900 (2022). PMLR
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Wankhade, M., Rao, A.C.S., Kulkarni, C.: A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review 55
2022
Cited alongside, same era.
Ji, J., Ma, Y., Sun, X., Zhou, Y., Wu, Y., Ji, R.: Knowing what to learn: a metric-oriented focal mechanism for image captioning. IEEE Transactions on Image Processing 31
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Liu, F., Emerson, G., Collier, N.: Visual spatial reasoning. Transactions of the Association for Computational Linguistics 11
2023
Later among the works it cites.
wendlerc: RenderedText. https://huggingface.co/datasets/wendlerc/RenderedText (2023)
2023
Later among the works it cites.
Kamizuru00: Diagram image-to-text. https://huggingface.co/datasets/Kamizuru00/diagram_image_to_text (2023)
2023
Later among the works it cites.
Tang, B., Boggust, A., Satyanarayan, A.: Vistext: A benchmark for semantically rich chart captioning (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Teknium: OpenHermes 2.5: An Open Dataset of Synthetic Data for Generalist LLM Assistants. HuggingFace (2023). https://huggingface.co/datasets/teknium/OpenHermes-2.5
2023
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., et al
2024
Later among the works it cites.
Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., Gao, J.: Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36
2024
Later among the works it cites.
2024
Later among the works it cites.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural information processing systems 36
2024
Later among the works it cites.
Liu, H., Li, C., Li, Y., Lee, Y.J.: Improved baselines with visual instruction tuning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26296–26306 (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Zhang, X., Yin, B.-W., Chen, Y., Lin, Z., Li, Y., Hou, Q., Cheng, M.-M.: Temo: Towards text-driven 3d stylization for multi-object meshes. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19531–19540 (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
Chen, D., Liu, J., Dai, W., Wang, B.: Visual instruction tuning with polite flamingo. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 17745–17753 (2024)
2024
Later among the works it cites.
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al.: Lima: Less is more for alignment. Advances in Neural Information Processing Systems 36
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Li*, B., Zhang*, P., Zhang*, K., Pu*, F., Du, X., Dong, Y., Liu, H., Zhang, Y., Zhang, G., Li, C., Liu, Z.: LMMs-Eval: Accelerating the Development of Large Multimoal Models. Zenodo (2024). https://github.com/EvolvingLMMs-Lab/lmms-eval
2024
Later among the works it cites.
Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., Wang, X., Chen, L., Huang, F., Yacoob, Y., et al
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
Zhang, X., Yin, B., Lin, Z., Hou, Q., Fan, D.-P., Cheng, M.-M.: Referring camouflaged object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)
2025
Closest in time.