Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have emerged as a central focus in both industry and academia, but often suffer from biases introduced by visual and language priors, which can lead to multimodal hallucination.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
A causally formulated hazard ratio estimation through backdoor adjustment on structural causal model
Riddhiman Adib, Paul Griffin, Sheikh Iqbal Ahamed, and Mohammad Adibuzzaman · 2020
Earlier work this paper cites.
Deep structural causal models for tractable counterfactual inference
Nick Pawlowski, Daniel Coelho de Castro, and Ben Glocker · 2020
Earlier work this paper cites.
Causality learning: A new perspective for interpretable machine learning
Guandong Xu, Tri Dung Duong, Qian Li, Shaowu Liu, and Xianzhi Wang · 2020
Earlier work this paper cites.
Counterfactual attention learning for fine-grained visual categorization and re-identification
Yongming Rao, Guangyi Chen, Jiwen Lu, and Jie Zhou · 2021
Earlier work this paper cites.
Causalvae: Disentangled representation learning via neural structural causal models
Mengyue Yang, Furui Liu, Zhitang Chen, Xinwei Shen, Jianye Hao, and Jun Wang · 2021
Earlier work this paper cites.
Rhino: Deep causal temporal relationship learning with history-dependent noise
Wenbo Gong, Joel Jennings, Cheng Zhang, and Nick Pawlowski · 2022
Earlier work this paper cites.
Modality, presentation, domain and training effects in statistical learning
Krisztina Sára Lukics and Ágnes Lukács · 2022
Earlier work this paper cites.
A-okvqa: A benchmark for visual question answering using world knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi · 2022
Earlier work this paper cites.
Predicting cellular responses with variational causal inference and refined relational information
Yulun Wu, Robert A Barton, Zichen Wang, Vassilis N Ioannidis, Carlo De Donno, Layne C Price, Luis F Voloch, and George Karypis · 2022
Earlier work this paper cites.
Overcoming language priors in vqa via adding visual module
Jia Zhao, Xuesong Zhang, Xuefeng Wang, Ying Yang, and Gang Sun · 2022
Earlier work this paper cites.
Cuts: Neural causal discovery from irregular time-series data
Yuxiao Cheng, Runzhao Yang, Tingxiong Xiao, Zongren Li, Jinli Suo, Kunlun He, and Qionghai Dai · 2023
Cited alongside, same era.
Dola: Decoding by contrasting layers improves factuality in large language models
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Cited alongside, same era.
Causal reasoning and large language models: Opening a new frontier for causality
Emre Kıcıman, Robert Ness, Amit Sharma, and Chenhao Tan · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Visual attention methods in deep learning: An in-depth survey
Mohammed Hassanin, Saeed Anwar, Ibrahim Radwan, Fahad Shahbaz Khan, and Ajmal Mian · 2024
Closest in time.
Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu · 2024
Closest in time.
Mmneuron: Discovering neuron-level domain-specific interpretation in multimodal large language model
Jiahao Huo, Yibo Yan, Boren Hu, Yutao Yue, and Xuming Hu · 2024
Closest in time.
Efficient multimodal large language models: A survey
Yizhang Jin, Jian Li, Yexin Liu, Tianjun Gu, Kai Wu, Zhengkai Jiang, Muyang He, Bo Zhao, Xin Tan, Zhenye Gan, et al · 2024
Closest in time.
Vlind-bench: Measuring language priors in large vision-language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen · 2023
Cited alongside, same era.
An empirical study on the language modal in visual question answering
Daowan Peng, Wei Wei, Xian-Ling Mao, Yuanyuan Fu, and Dangyang Chen · 2023
Cited alongside, same era.
Causal inference using llm-guided discovery
Aniket Vashishtha, Abbavaram Gowtham Reddy, Abhinav Kumar, Saketh Bachu, Vineeth N Balasubramanian, and Amit Sharma · 2023
Cited alongside, same era.
Language in a bottle: Language model guided concept bottlenecks for interpretable image classification
Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, and Mark Yatskar · 2023
Cited alongside, same era.
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen · 2023
Cited alongside, same era.
Dpnet: Dynamic poly-attention network for trustworthy multi-modal classification
Xin Zou, Chang Tang, Xiao Zheng, Zhenglai Li, Xiao He, Shan An, and Xinwang Liu · 2023
Cited alongside, same era.
Hallucination of multimodal large language models: A survey
Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou · 2024
Cited alongside, same era.
Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective
Meiqi Chen, Yixin Cao, Yan Zhang, and Chaochao Lu · 2024
Cited alongside, same era.
Kang-il Lee, Minbeom Kim, Seunghyun Yoon, Minsung Kim, Dongryeol Lee, Hyukhun Koh, and Kyomin Jung · 2024
Closest in time.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin · 2024
Closest in time.
Georeasoner: Reasoning on geospatially grounded context for natural language understanding
Yibo Yan and Joey Lee · 2024
Closest in time.
Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web
Yibo Yan, Haomin Wen, Siru Zhong, Wei Chen, Haodong Chen, Qingsong Wen, Roger Zimmermann, and Yuxuan Liang · 2024
Closest in time.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, and Fei Huang · 2024
Closest in time.
Mm-llms: Recent advances in multimodal large language models
Duzhen Zhang, Yahan Yu, Chenxing Li, Jiahua Dong, Dan Su, Chenhui Chu, and Dong Yu · 2024
Closest in time.
Enhancing contextual understanding in large language models through contrastive decoding
Zheng Zhao, Emilio Monti, Jens Lehmann, and Haytham Assem · 2024
Closest in time.
Kening Zheng, Junkai Chen, Yibo Yan, Xin Zou, and Xuming Hu · 2024
Closest in time.