Fetching the paper…
Reading the bibliography…
Multimodal foundation models are prone to hallucination, generating outputs that either contradict the input or are not grounded by factual information.
Evaluations of self and others: Self-enhancement biases in social judgments
J. D. Brown · 1986
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
A. See, P. J. Liu, and C. D. Manning · 2017
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Earlier work this paper cites.
Object hallucination in image captioning
A. Rohrbach, L. A. Hendricks, K. Burns, T. Darrell, and K. Saenko · 2018
Earlier work this paper cites.
How2: A large-scale dataset for multimodal language understanding
R. Sanabria, O. Caglayan, S. Palaskar, D. Elliott, L. Barrault, L. Specia, and F. Metze · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal · 2018
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
E. Dinan, S. Roller, K. Shuster, A. Fan, M. Auli, and J. Weston · 2019
Earlier work this paper cites.
Episodic memory reader: Learning what to remember for question answering from streaming data
M. Han, M. Kang, H. Jung, and S. J. Hwang · 2019
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
J. Maynez, S. Narayan, B. Bohnet, and R. McDonald · 2020
Earlier work this paper cites.
NExT-QA: Next phase of question-answering to explaining temporal actions
J. Xiao, X. Shang, A. Yao, and T.-S. Chua · 2021
Earlier work this paper cites.
Faithdial: A faithful benchmark for information-seeking dialogue
N. Dziri, E. Kamalloo, S. Milton, O. Zaiane, M. Yu, E. Ponti, and S. Reddy · 2022
Earlier work this paper cites.
Learning to answer questions in dynamic audio-visual scenarios
G. li, Y. Wei, Y. Tian, C. Xu, J.-R. Wen, and D. Hu · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
The falcon series of open language models
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah, Étienne Goffinet, D. Hesslow, J. Launay, Q. Malartic, D. Mazzotta, B. Noune, B. Pannier, and G. Penedo · 2023
Earlier work this paper cites.
Valor: Vision-audio-language omni-perception pretraining model and dataset
S. Chen, X. He, L. Guo, X. Zhu, W. Wang, J. Tang, and J. Liu · 2023
Earlier work this paper cites.
Vicuna: An opensource chatbot impressing gpt-4 with 90% chatgpt quality., 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Earlier work this paper cites.
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models
Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Cited alongside, same era.
FactKB: Generalizable factuality evaluation using language models enhanced with factual knowledge
S. Feng, V. Balachandran, Y. Bai, and Y. Tsvetkov · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2023
Cited alongside, same era.
HaluEval: A large-scale hallucination evaluation benchmark for large language models
J. Li, X. Cheng, X. Zhao, J.-Y. Nie, and J.-R. Wen · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Y. Li, Y. Du, K. Zhou, J. Wang, X. Zhao, and J.-R. Wen · 2023
Cited alongside, same era.
A survey on detection of llms-generated content
X. Yang, L. Pan, X. Zhao, H. Chen, L. Petzold, W. Y. Wang, and W. Cheng · 2023
Later among the works it cites.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, F. Huang, and J. Zhou · 2023
Later among the works it cites.
Video-LLaMA: An instruction-tuned audio-visual language model for video understanding
H. Zhang, X. Li, and L. Bing · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023
B. Zhu, E. Frick, T. Wu, H. Zhu, and J. Jiao · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Video-llava: Learning united visual representation by alignment before projection
B. Lin, B. Zhu, Y. Ye, M. Ning, P. Jin, and L. Yuan · 2023
Cited alongside, same era.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
Valley: Video assistant with large language model enhanced ability
R. Luo, Z. Zhao, M. Yang, J. Dong, M. Qiu, P. Lu, T. Wang, and Z. Wei · 2023
Cited alongside, same era.
SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models
P. Manakul, A. Liusie, and M. Gales · 2023
Cited alongside, same era.
Inverse scaling: When bigger isn’t better
I. R. McKenzie, A. Lyzhov, M. Pieler, A. Parrish, A. Mueller, A. Prabhu, E. McLean, A. Kirtland, A. Ross, A. Liu, A. Gritsevskiy, D. Wurgaft, D. Kauffman, G. Recchia, J. Liu, J. Cavanagh, M. Weiss, S. Huang, T. F. Droid, T. Tseng, T. Korbak, X. Shen, Y. Zhang, Z. Zhou, N. Kim, S. R. Bowman, and E. Perez · 2023
Cited alongside, same era.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi · 2023
Cited alongside, same era.
Orca: Progressive learning from complex explanation traces of gpt-4
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, and A. Awadallah · 2023
Cited alongside, same era.
Later among the works it cites.
Listen, think, and understand
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass · 2024
Closest in time.
Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, D. Manocha, and T. Zhou · 2024
Closest in time.
Chat-univi: Unified visual representation empowers large language models with image and video understanding
P. Jin, R. Takanobu, C. Zhang, X. Cao, and L. Yuan · 2024
Closest in time.
A survey on hallucination in large vision-language models
H. Liu, W. Xue, Y. Chen, D. Chen, X. Zhao, K. Wang, L. Hou, R. Li, and W. Peng · 2024
Closest in time.
M. Nahar, H. Seo, E.-J. Lee, A. Xiong, and D. Lee · 2024
Closest in time.
SALMONN: Towards generic hearing abilities for large language models
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. Team · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
S. M. T. I. Tonmoy, S. M. M. Zaman, V. Jain, A. Rani, V. Rawte, A. Chadha, and A. Das · 2024
Closest in time.
Replacing judges with juries: Evaluating llm generations with a panel of diverse models
P. Verga, S. Hofstatter, S. Althammer, Y. Su, A. Piktus, A. Arkhangorodsky, M. Xu, N. White, and P. Lewis · 2024
Closest in time.
Measuring and reducing llm hallucination without gold-standard answers via expertise-weighting
J. Wei, Y. Yao, J.-F. Ton, H. Guo, A. Estornell, and Y. Liu · 2024
Closest in time.
Analyzing and mitigating object hallucination in large vision-language models
Y. Zhou, C. Cui, J. Yoon, L. Zhang, Z. Deng, C. Finn, M. Bansal, and H. Yao · 2024
Closest in time.