Fetching the paper…
Reading the bibliography…
Although Large Visual Language Models (LVLMs) have demonstrated exceptional abilities in understanding multimodal data, they invariably suffer from hallucinations, leading to a disconnect between the generated text and the corresponding images.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Audio chord recognition with recurrent neural networks
N. Boulanger-Lewandowski, Y. Bengio, and P. Vincent · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Hierarchical neural story generation
A. Fan, M. Lewis, and Y. Dauphin · 2018
Earlier work this paper cites.
Hallucinations in neural machine translation
K. Lee, O. Firat, A. Agarwal, C. Fannjiang, and D. Sussillo · 2018
Earlier work this paper cites.
Object hallucination in image captioning
A. Rohrbach, L. A. Hendricks, K. Burns, T. Darrell, and K. Saenko · 2018
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Earlier work this paper cites.
Detecting hallucinated content in conditional neural sequence generation
C. Zhou, G. Neubig, J. Gu, M. Diab, P. Guzman, L. Zettlemoyer, and M. Ghazvininejad · 2020
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
X. L. Li, A. Holtzman, D. Fried, P. Liang, J. Eisner, T. Hashimoto, L. Zettlemoyer, and M. Lewis · 2022
Earlier work this paper cites.
Natural color fool: Towards boosting black-box unrestricted attacks
S. Yuan, Q. Zhang, L. Gao, Y. Cheng, and J. Song · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Dola: Decoding by contrasting layers improves factuality in large language models
Y.-S. Chuang, Y. Xie, H. Luo, Y. Kim, J. Glass, and P. He · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, Z. Qiu, W. Lin, J. Yang, X. Zheng, et al · 2023
Cited alongside, same era.
Woodpecker: Hallucination correction for multimodal large language models
S. Yin, C. Fu, S. Zhao, T. Xu, H. Wang, D. Sui, Y. Shen, K. Li, X. Sun, and E. Chen · 2023
Later among the works it cites.
T. Yu, Y. Yao, H. Zhang, T. He, Y. Han, G. Cui, J. Hu, Z. Liu, H.-T. Zheng, M. Sun, et al · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen, et al · 2023
Later among the works it cites.
Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization
Z. Zhao, B. Wang, L. Ouyang, X. Dong, J. Wang, and C. He · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Gao, J. Schulman, and J. Hilton · 2023
Cited alongside, same era.
Detecting and preventing hallucinations in large vision language models
A. Gunjal, J. Yin, and E. Bas · 2023
Cited alongside, same era.
Q. Huang, X. Dong, P. Zhang, B. Wang, C. He, J. Wang, D. Lin, W. Zhang, and N. Yu · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung · 2023
Cited alongside, same era.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
S. Leng, H. Zhang, G. Chen, X. Li, S. Lu, C. Miao, and L. Bing · 2023
Cited alongside, same era.
Negative object presence evaluation (nope) to measure object hallucination in vision-language models
H. Lovenia, W. Dai, S. Cahyawijaya, Z. Ji, and P. Fung · 2023
Cited alongside, same era.
MathVista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Cited alongside, same era.
GPT-4V(ision) system card
OpenAI · 2023
Cited alongside, same era.
Y. Zhou, C. Cui, J. Yoon, L. Zhang, Z. Deng, C. Finn, M. Bansal, and H. Yao · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Halc: Object hallucination reduction via adaptive focal-contrast decoding
Z. Chen, Z. Zhao, H. Luo, H. Yao, B. Li, and J. Zhou · 2024
Closest in time.
Multi-modal hallucination control by visual information grounding
A. Favero, L. Zancato, M. Trager, S. Choudhary, P. Perera, A. Achille, A. Swaminathan, and S. Soatto · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Any target can be offense: Adversarial example generation via generalized latent infection
Y. Sun, S. Yuan, X. Wang, L. Gao, and J. Song · 2024
Closest in time.
Mitigating hallucinations in large vision-language models with instruction contrastive decoding
X. Wang, J. Pan, L. Ding, and C. Biemann · 2024
Closest in time.
Debiasing large visual language models
Y.-F. Zhang, W. Yu, Q. Wen, X. Wang, Z. Zhang, L. Wang, R. Jin, and T. Tan · 2024
Closest in time.
Aligning modalities in vision large language models via preference fine-tuning
Y. Zhou, C. Cui, R. Rafailov, C. Finn, and H. Yao · 2024
Closest in time.
Ibd: Alleviating hallucinations in large vision-language models via image-biased decoding
L. Zhu, D. Ji, T. Chen, P. Xu, J. Ye, and J. Liu · 2024
Closest in time.