Fetching the paper…
Reading the bibliography…
Despite achieving rapid developments and with widespread applications, Large Vision-Language Models (LVLMs) confront a serious challenge of being prone to generating hallucinations.
Audio chord recognition with recurrent neural networks
Boulanger-Lewandowski, N., Bengio, Y., and Vincent, P · 2013
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
Teerapittayanon, S., McDanel, B., and Kung, H.-T · 2016
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Honnibal, M. and Montani, I · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Earlier work this paper cites.
Object hallucination in image captioning
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K · 2018
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Visual perturbation-aware collaborative learning for overcoming the language prior problem
Han, Y., Nie, L., Yin, J., Wu, J., and Yan, Y · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Mocha: Multi-objective reinforcement mitigating caption hallucinations
Ben-Kish, A., Yanuka, M., Alper, M., Giryes, R., and Averbuch-Elor, H · 2023
Cited alongside, same era.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R · 2023
Cited alongside, same era.
Dola: Decoding by contrasting layers improves factuality in large language models
Chuang, Y.-S., Xie, Y., Luo, H., Kim, Y., Glass, J., and He, P · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Vary: Scaling up the vision vocabulary for large vision-language models
Wei, H., Kong, L., Chen, J., Zhao, L., Ge, Z., Yang, J., Sun, J., Han, C., and Zhang, X · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N · 2023
Later among the works it cites.
Xu, J., Zhou, X., Yan, S., Gu, X., Arnab, A., Sun, C., Wang, X., and Schmid, C · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al · 2023
Cited alongside, same era.
Huang, Q., Dong, X., Zhang, P., Wang, B., He, C., Wang, J., Lin, D., Zhang, W., and Yu, N · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P · 2023
Cited alongside, same era.
Exposing and mitigating spurious correlations for cross-modal retrieval
Kim, J. M., Koepke, A., Schmid, C., and Akata, Z · 2023
Cited alongside, same era.
Lisa: Reasoning segmentation via large language model
Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J · 2023
Cited alongside, same era.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., and Bing, L · 2023
Cited alongside, same era.
Glamm: Pixel grounding large multimodal model
Rasheed, H., Maaz, M., Shaji, S., Shaker, A., Khan, S., Cholakkal, H., Anwer, R. M., Xing, E., Yang, M.-H., and Khan, F. S · 2023
Cited alongside, same era.
Li, J., Li, D., Savarese, S., and Hoi, S
Cited in the paper.
Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., and Chen, E · 2023
Later among the works it cites.
Ferret: Refer and ground anything anywhere at any granularity, 2023
You, H., Zhang, H., Gan, Z., Du, X., Zhang, B., Wang, Z., Cao, L., Chang, S.-F., and Yang, Y · 2023
Later among the works it cites.
Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data
Yu, Q., Li, J., Wei, L., Pang, L., Ye, W., Qin, B., Tang, S., Tian, Q., and Zhuang, Y · 2023
Later among the works it cites.
Pmc-vqa: Visual instruction tuning for medical visual question answering
Zhang, X., Wu, C., Zhao, Z., Lin, W., Zhang, Y., Wang, Y., and Xie, W · 2023
Later among the works it cites.
Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization
Zhao, Z., Wang, B., Ouyang, L., Dong, X., Wang, J., and He, C · 2023
Later among the works it cites.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H · 2023
Later among the works it cites.