Fetching the paper…
Reading the bibliography…
The rapid advancement of foundation models (FMs) across language, image, audio, and video domains has shown remarkable capabilities in diverse tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Video-based facial expression hallucination: A two-level hierarchical fusion approach
Jian Zhang, Yueting Zhuang, and Fei Wu. 2006 · 2006
Earlier work this paper cites.
Content matters: A qualitative analysis of verbal hallucinations
Nienke Moernaut, Isabella Leudar, and Thomas Verdooren. 2018 · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason Corso. 2018 · 2018
Earlier work this paper cites.
Parallelized convolutional recurrent neural network with spectral features for speech emotion recognition
Pengxu Jiang, Hongliang Fu, Huawei Tao, Peizhi Lei, and Li Zhao. 2019 · 2019
Earlier work this paper cites.
Streamlined dense video captioning
Jonghwan Mun, Linjie Yang, Zhou Ren, Ning Xu, and Bohyung Han. 2019 · 2019
Earlier work this paper cites.
Multi-modal dense video captioning
Vladimir Iashin and Esa Rahtu. 2020 · 2020
Earlier work this paper cites.
An efficient framework for dense video captioning
Maitreya Suin and AN Rajagopalan. 2020 · 2020
Earlier work this paper cites.
Automated audio captioning with weakly supervised pre-training and word selection methods
Qichen Han, Weiqiang Yuan, Dong Liu, Xiang Li, and Zhen Yang. 2021 · 2021
Earlier work this paper cites.
Internal video inpainting by implicit long-range propagation
Hao Ouyang, Tengfei Wang, and Qifeng Chen. 2021 · 2021
Earlier work this paper cites.
Zhongjie Ye, Helin Wang, Dongchao Yang, and Yuexian Zou. 2021 · 2021
Earlier work this paper cites.
Maxm: Towards multilingual visual question answering
Soravit Changpinyo, Linting Xue, Idan Szpektor, Ashish V Thapliyal, Julien Amelot, Michal Yarom, Xi Chen, and Radu Soricut. 2022 · 2022
Earlier work this paper cites.
Plausible may not be faithful: Probing object hallucination in vision-language pre-training
Wenliang Dai, Zihan Liu, Ziwei Ji, Dan Su, and Pascale Fung. 2022 · 2022
Earlier work this paper cites.
Rarr: Researching and revising what language models say, using language models
Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al. 2022 · 2022
Earlier work this paper cites.
Multi-granularity aggregation transformer for joint video-audio-text representation learning
Mengge He, Wenjing Du, Zhiquan Wen, Qing Du, Yutong Xie, and Qi Wu. 2022 · 2022
Earlier work this paper cites.
Diffusion models for video prediction and infilling
Tobias Höppe, Arash Mehrjou, Stefan Bauer, Didrik Nielsen, and Andrea Dittadi. 2022 · 2022
Earlier work this paper cites.
Valhalla: Visual hallucination for machine translation. in 2022 ieee
Y Li, R Panda, Y Kim, C Chen, R Feris, D Cox, and N Vasconcelos. 2022 · 2022
Earlier work this paper cites.
Emscore: Evaluating video captioning via coarse-grained and fine-grained embedding matching
Yaya Shi, Xu Yang, Haiyang Xu, Chunfeng Yuan, Bing Li, Weiming Hu, and Zheng-Jun Zha. 2022 · 2022
Earlier work this paper cites.
Adversarial semantic hallucination for domain generalized semantic segmentation
Gabriel Tjio, Ping Liu, Joey Tianyi Zhou, and Rick Siow Mong Goh. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar. 2022 · 2022
Earlier work this paper cites.
Creating trustworthy llms: Dealing with hallucinations in healthcare ai
Muhammad Aurangzeb Ahmad, Ilker Yaramis, and Taposh Dutta Roy. 2023 · 2023
Earlier work this paper cites.
The internal state of an llm knows when its lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Earlier work this paper cites.
Touchstone: Evaluating vision-language models by language models
Shuai Bai, Shusheng Yang, Jinze Bai, Peng Wang, Xingxuan Zhang, Junyang Lin, Xinggang Wang, Chang Zhou, and Jingren Zhou. 2023 · 2023
Earlier work this paper cites.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2023 · 2023
Earlier work this paper cites.
Evaluating hallucinations in chinese large language models
Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang, Siyin Wang, Xiangyang Liu, Mozhi Zhang, Junliang He, Mianqiu Huang, Zhangyue Yin, Kai Chen, et al. 2023 · 2023
Earlier work this paper cites.
I Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, Pengfei Liu, et al. 2023 · 2023
Earlier work this paper cites.
Clearvid: Curriculum learning for video description
Cheng-Yu Chuang and Pooyan Fazli. 2023 · 2023
Earlier work this paper cites.
Dola: Decoding by contrasting layers improves factuality in large language models
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. 2023 · 2023
Earlier work this paper cites.
Chatlaw: Open-source legal large language model with integrated external knowledge bases
Jiaxi Cui, Zongjian Li, Yang Yan, Bohua Chen, and Li Yuan. 2023 · 2023
Earlier work this paper cites.
Chain-of-verification reduces hallucination in large language models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023 · 2023
Earlier work this paper cites.
Lp-musiccaps: Llm-based pseudo music captioning
SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam. 2023 · 2023
Cited alongside, same era.
Halo: Estimation and reduction of hallucinations in open-source weak large language models
Mohamed Elaraby, Mengyin Lu, Jacob Dunn, Xueying Zhang, Yu Wang, and Shizhu Liu. 2023 · 2023
Cited alongside, same era.
Text-to-audio generation using instruction-tuned llm and latent diffusion model
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria. 2023 · 2023
Cited alongside, same era.
Compa: Addressing the gap in compositional reasoning in audio-language models
Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Evuru, S Ramaneswaran, S Sakshi, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha. 2023 · 2023
Cited alongside, same era.
Explaining legal concepts with augmented large language models (gpt-4)
Jaromir Savelka, Kevin D Ashley, Morgan A Gray, Hannes Westermann, and Huihui Xu. 2023 · 2023
Later among the works it cites.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian. 2023 · 2023
Later among the works it cites.
Aligning large multimodal models with factually augmented rlhf
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. 2023 · 2023
Later among the works it cites.
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, et al. 2023 · 2023
Cited alongside, same era.
Hallucinations in large multilingual translation models
Nuno M Guerreiro, Duarte M Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, and André FT Martins. 2023 · 2023
Cited alongside, same era.
Let’s think frame by frame with vip: A video infilling and prediction dataset for evaluating video chain-of-thought
Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He, Alex Mei, Yujie Lu, Chinmay Sonar, Michael Saxon, and William Wang. 2023 · 2023
Cited alongside, same era.
Ciem: Contrastive instruction evaluation method for better instruction tuning
Hongyu Hu, Jiyuan Zhang, Minyi Zhao, and Zhenbang Sun. 2023 · 2023
Cited alongside, same era.
Citation: A key to building responsible and accountable large language models
Jie Huang and Kevin Chen-Chuan Chang. 2023 · 2023
Cited alongside, same era.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023 · 2023
Cited alongside, same era.
Dehallucinating large language models using formal methods guided iterative prompting
Susmit Jha, Sumit Kumar Jha, Patrick Lincoln, Nathaniel D Bastian, Alvaro Velasquez, and Sandeep Neema. 2023 · 2023
Cited alongside, same era.
Towards mitigating llm hallucination via self reflection
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
Evaluation and analysis of hallucination in large vision-language models
Junyang Wang, Yiyang Zhou, Guohai Xu, Pengcheng Shi, Chenlin Zhao, Haiyang Xu, Qinghao Ye, Ming Yan, Ji Zhang, Jihua Zhu, et al. 2023 · 2023
Later among the works it cites.
A context-aware model with a pre-trained context encoder for dense video captioning
Weilun Wu and Yang Gao. 2023 · 2023
Later among the works it cites.
A new benchmark and reverse validation method for passage-level hallucination detection
Shiping Yang, Renliang Sun, and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization
Zhiyuan Zhao, Bin Wang, Linke Ouyang, Xiaoyi Dong, Jiaqi Wang, and Conghui He. 2023 · 2023
Later among the works it cites.
Analyzing and mitigating object hallucination in large vision-language models
Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.
Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies
Ke Chen, Yusong Wu, Haohe Liu, Marianna Nezhurina, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2024a · 2024
Closest in time.
Large legal fictions: Profiling legal hallucinations in large language models
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho. 2024 · 2024
Closest in time.
Ailin Deng, Zhirui Chen, and Bryan Hooi. 2024 · 2024
Closest in time.
Pam: Prompting audio-language models for audio quality assessment
Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde, Hannes Gamper, Mahmoud Al Ismail, Rita Singh, Bhiksha Raj, and Huaming Wang. 2024 · 2024
Closest in time.
Natural language supervision for general-purpose audio representations
Benjamin Elizalde, Soham Deshmukh, and Huaming Wang. 2024 · 2024
Closest in time.
Recap: retrieval-augmented audio captioning
Sreyan Ghosh, Sonal Kumar, Chandra Kiran Reddy Evuru, Ramani Duraiswami, and Dinesh Manocha. 2024e · 2024
Closest in time.
Detecting and preventing hallucinations in large vision language models
Anisha Gunjal, Jihan Yin, and Erhan Bas. 2024 · 2024
Closest in time.
Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning
Jaeyeon Kim, Jaeyoon Jung, Jinjoo Lee, and Sang Hoon Woo. 2024 · 2024
Closest in time.
Evaluation and enhancement of semantic grounding in large vision-language models
Jiaying Lu, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo, Yawen Zhang, Baochen Sun, Carl Yang, and Jie Yang. 2024 · 2024
Closest in time.
On the audio hallucinations in large audio-video language models
Taichi Nishimura, Shota Nakada, and Masayoshi Kondo. 2024 · 2024
Closest in time.
Journey of hallucination-minimized generative ai solutions for financial decision makers
Sohini Roychowdhury. 2024 · 2024
Closest in time.
Enhancing adverse drug event detection with multimodal dataset: Corpus creation and model development
Pranab Sahoo, Ayush Singh, Sriparna Saha, Aman Chadha, and Samrat Mondal. 2024a · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. 2024 · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
SM Tonmoy, SM Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024 · 2024
Closest in time.
Vigor: Improving visual grounding of large vision language models with fine-grained reward modeling
Siming Yan, Min Bai, Weifeng Chen, Xiong Zhou, Qixing Huang, and Li Erran Li. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Retrieval-augmented text-to-audio generation
Yi Yuan, Haohe Liu, Xubo Liu, Qiushi Huang, Mark D Plumbley, and Wenwu Wang. 2024 · 2024
Closest in time.
Mitigating object hallucination in large vision-language models via classifier-free guidance
Linxi Zhao, Yihe Deng, Weitong Zhang, and Quanquan Gu. 2024 · 2024
Closest in time.
Streaming dense video captioning
Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, and Cordelia Schmid. 2024 · 2024
Closest in time.
Cacophony: An improved contrastive audio-text model
Ge Zhu and Zhiyao Duan. 2024 · 2024
Closest in time.