Fetching the paper…
Reading the bibliography…
Hallucination has been a major problem for large language models and remains a critical challenge when it comes to multimodality in which vision-language models (VLMs) have to deal with not just textual but also visual inputs.
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: Bleu: a method for automatic evaluation of machine translation (10 2002). https://doi.org/10.3115/1073083.1073135
2002
Earlier work this paper cites.
Lin, C.Y.: ROUGE: A package for automatic evaluation of summaries. In: Text Summarization Branches Out. pp. 74–81. Association for Computational Linguistics, Barcelona, Spain (Jul 2004), https://aclanthology.org/W04-1013
2004
Earlier work this paper cites.
Banerjee, S., Lavie, A.: METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In: Goldstein, J., Lavie, A., Lin, C.Y., Voss, C. (eds.) Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization. pp. 65–72. Association for Computational Linguistics, Ann Arbor, Michigan (Jun 2005), https://aclanthology.org/W05-0909
2005
Earlier work this paper cites.
Sedgwick, P.: Pearson’s correlation coefficient. BMJ 345
2012
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13. pp. 740–755. Springer (2014)
2014
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: Vqa: Visual question answering. In: Proceedings of the IEEE international conference on computer vision. pp. 2425–2433 (2015)
2015
Earlier work this paper cites.
Vedantam, R., Lawrence Zitnick, C., Parikh, D.: Cider: Consensus-based image description evaluation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4566–4575 (2015)
2015
Earlier work this paper cites.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 6904–6913 (2017)
2017
Earlier work this paper cites.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al.: Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision 123
2017
Earlier work this paper cites.
Rohrbach, A., Hendricks, L.A., Burns, K., Darrell, T., Saenko, K.: Object hallucination in image captioning. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. pp. 4035–4045 (2018)
2018
Earlier work this paper cites.
Marino, K., Rastegari, M., Farhadi, A., Mottaghi, R.: Ok-vqa: A visual question answering benchmark requiring external knowledge. In: Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Earlier work this paper cites.
Zellers, R., Bisk, Y., Farhadi, A., Choi, Y.: From recognition to cognition: Visual commonsense reasoning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 6720–6731 (2019)
2019
Earlier work this paper cites.
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., et al.: The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. International Journal of Computer Vision 128
2020
Earlier work this paper cites.
Bender, E.M., Gebru, T., McMillan-Major, A., Shmitchell, S.: On the dangers of stochastic parrots: Can language models be too big? In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. pp. 610–623 (2021)
2021
Earlier work this paper cites.
Biten, A.F., Gómez, L., Karatzas, D.: Let there be a clock on the beach: Reducing object hallucination in image captioning. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. pp. 1381–1390 (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10684–10695 (2022)
2022
Earlier work this paper cites.
Schwenk, D., Khandelwal, A., Clark, C., Marino, K., Mottaghi, R.: A-okvqa: A benchmark for visual question answering using world knowledge. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part VIII. pp. 146–162. Springer (2022)
2022
Earlier work this paper cites.
Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., Yang, H.: Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In: International Conference on Machine Learning. pp. 23318–23340. PMLR (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Gpt-4v(ision) system card (2023), https://api.semanticscholar.org/CorpusID:263218031
2023
Cited alongside, same era.
Midjourney. https://midjourney.com/ (2023)
2023
Cited alongside, same era.
Alkaissi, H., Mcfarlane, S.: Artificial hallucinations in chatgpt: Implications in scientific writing. Cureus 15
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Peng, D.: Langfun (Sep 2023), https://github.com/google/langfun
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bitton-Guetta, N., Bitton, Y., Hessel, J., Schmidt, L., Elovici, Y., Stanovsky, G., Schwartz, R.: Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional images. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 2616–2627 (2023)
2023
Cited alongside, same era.
Chiang, W.L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J.E., et al.: Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023) (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A., Fung, P.: Survey of hallucination in natural language generation. ACM Computing Surveys 55
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
Wang, H., Wu, W., Dou, Z., He, L., Yang, L.: Performance and exploration of chatgpt in medical examination, records and education in chinese: Pave the way for medical ai. International Journal of Medical Informatics 177
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural information processing systems 36
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Zhang, Q., Zhang, J., Xu, Y., Tao, D.: Vision transformer with quadrangle attention. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
Closest in time.