Fetching the paper…
Reading the bibliography…
While Large Vision-Language Models (LVLMs) have rapidly advanced in recent years, the prevalent issue known as the `hallucination' problem has emerged as a significant bottleneck, hindering their real-world deployments.
Efficient wavelet boost learning-based multi-stage progressive refinement network for underwater image enhancement
Fushuo Huo, Bingheng Li, and Xuegui Zhu · 1952
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, and et al · 2014
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, and et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, and et al · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, and et al · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, and et al · 2021
Earlier work this paper cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, and et al · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, and et al · 2022
Earlier work this paper cites.
Factuality enhanced language models for open-ended text generation
Nayeon Lee, Wei Ping, and et al · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, and et al · 2022
Earlier work this paper cites.
Confident adaptive language modeling
Tal Schuster, Adam Fisch, and et al · 2022
Earlier work this paper cites.
A-okvqa: A benchmark for visual question answering using world knowledge
Dustin Schwenk, Apoorv Khandelwal, and et al · 2022
Earlier work this paper cites.
Introducing our multimodal models, 2023
Rohan Bavishi, Erich Elsen, and et al · 2023
Earlier work this paper cites.
Token merging: Your ViT but faster
Daniel Bolya, Cheng-Yang Fu, and et al · 2023
Earlier work this paper cites.
PuMer: Pruning and merging tokens for efficient vision language models
Qingqing Cao, Bhargavi Paranjape, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang and Zhuohan et al Li · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai and Junnan Li et al · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, and et al · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, and et al · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Tianrui Guan, Fuxiao Liu, and et al · 2024
Closest in time.
Detecting and preventing hallucinations in large vision language models
Anisha Gunjal, Jihan Yin, and Erhan Bas · 2024
Closest in time.
Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu · 2024
Closest in time.
Hallucination augmented contrastive learning for multimodal large language model
Chaoya Jiang, Haiyang Xu, and et al · 2024
Closest in time.
Instructive decoding: Instruction-tuned large language models are self-refiner from noisy instructions
Taehyeon Kim, Joonkee Kim, Gihun Lee, and Se-Young Yun · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Potsawee Manakul, Adian Liusie, and Mark JF Gales · 2023
Cited alongside, same era.
Scaling open-vocabulary object detection
Matthias Minderer, Alexey Gritsenko, and Neil Houlsby · 2023
Cited alongside, same era.
Stanford alpaca: an instruction-following llama model (2023)
Rohan Taori, Ishaan Gulrajani, and et al · 2023
Cited alongside, same era.
Chatcad: Interactive computer-aided diagnosis on medical image using large language models
Sheng Wang, Zihao Zhao, and et al · 2023
Cited alongside, same era.
mplug-owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, and et al · 2023
Cited alongside, same era.
Woodpecker: Hallucination correction for multimodal large language models
Shukang Yin, Chaoyou Fu, and et al · 2023
Cited alongside, same era.
Gpt4roi: Instruction tuning large language model on region-of-interest
Shilong Zhang, Peize Sun, and et al · 2023
Cited alongside, same era.
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, and et al · 2024
Closest in time.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Sicong Leng, Hang Zhang, and et al · 2024
Closest in time.
Reqa: Coarse-to-fine assessment of image quality to alleviate the range effect
Bingheng Li and Fushuo Huo · 2024
Closest in time.
Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
Bo Li, Kaichen Zhang, and et al · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, and et al · 2024
Closest in time.
Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens
Fan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian, Jiashi Feng, and Yi Yang · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
AI Meta · 2024
Closest in time.
Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Yuzhang Shang, Mu Cai, and et al · 2024
Closest in time.
Don’t miss the forest for the trees: Attentional vision calibration for large vision language models
Sangmin Woo, Donguk Kim, and et al · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, and et al · 2024
Closest in time.
How language model hallucinations can snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith · 2024
Closest in time.
Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization
Zhiyuan Zhao, Bin Wang, and et al · 2024
Closest in time.
Analyzing and mitigating object hallucination in large vision-language models
Yiyang Zhou, Chenhang Cui, and et al · 2024
Closest in time.
Ibd: Alleviating hallucinations in large vision-language models via image-biased decoding
Lanyun Zhu, Deyi Ji, Tianrun Chen, Peng Xu, Jieping Ye, and Jun Liu · 2024
Closest in time.