Fetching the paper…
Reading the bibliography…
With the development of artificial intelligence, large-scale models have become increasingly intelligent.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2005
Earlier work this paper cites.
Overfitting or underfitting? understand robustness drop in adversarial training
Zichao Li, Liyuan Liu, Chengyu Dong, and Jingbo Shang. 2020 · 2010
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2014 · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Francis Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda B. Viégas, and Martin Wattenberg. 2017 · 2017
Earlier work this paper cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018 · 2018
Earlier work this paper cites.
Towards the unification and robustness of perturbation and gradient based explanations
Sushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay, Zhiwei Steven Wu, and Himabindu Lakkaraju. 2021 · 2021
Earlier work this paper cites.
Multi-object tracking with hallucinated and unlabeled videos
Daniel McKee, Bing Shuai, Andrew G. Berneshawi, Manchen Wang, Davide Modolo, Svetlana Lazebnik, and Joseph Tighe. 2021 · 2021
Earlier work this paper cites.
Multimodal entity tagging with multimodal knowledge base
Hao Peng, Hang Li, Lei Hou, Juanzi Li, and Chao Qiao. 2021 · 2021
Earlier work this paper cites.
Towards out-of-distribution generalization: A survey
Zheyan Shen, Jiashuo Liu, Yue He, Xingxuan Zhang, Renzhe Xu, Han Yu, and Peng Cui. 2021 · 2021
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Earlier work this paper cites.
Boosting entity-aware image captioning with multi-modal knowledge graph
Wentian Zhao, Yao Hu, Heda Wang, Xinxiao Wu, and Jiebo Luo. 2021 · 2021
Earlier work this paper cites.
On the origin of hallucinations in conversational models: Is it the datasets or the models?
Nouha Dziri, Sivan Milton, Mo Yu, Osmar R Zaiane, and Siva Reddy. 2022 · 2022
Earlier work this paper cites.
Posetriplet: Co-evolving 3d human pose estimation, imitation, and hallucination under self-supervision
Kehong Gong, Bingbing Li, Jianfeng Zhang, Tao Wang, Jing Huang, Michael Bi Mi, Jiashi Feng, and Xinchao Wang. 2022 · 2022
Earlier work this paper cites.
Training data influence analysis and estimation: A survey
Zayd Hammoudeh and Daniel Lowd. 2022 · 2022
Earlier work this paper cites.
Certified policy smoothing for cooperative multi-agent reinforcement learning
Ronghui Mu, Wenjie Ruan, Leandro Soriano Marcolino, Gaojie Jin, and Qiang Ni. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022 · 2022
Earlier work this paper cites.
Iterative teaching by data hallucination
Zeju Qiu, Weiyang Liu, Tim Z. Xiao, Zhen Liu, Umang Bhatt, Yucen Luo, Adrian Weller, and Bernhard Scholkopf. 2022 · 2022
Earlier work this paper cites.
Hallucination of speech recognition errors with sequence to sequence learning
Prashant Serai, Vishal Sunder, and Eric Fosler-Lussier. 2022 · 2022
Earlier work this paper cites.
Thinking hallucination for video captioning
Nasib Ullah and Partha Pratim Mohanta. 2022 · 2022
Earlier work this paper cites.
Information-theoretic text hallucination reduction for video-grounded dialogue
Sunjae Yoon, Eunseop Yoon, Hee Suk Yoon, Junyeong Kim, and Changdong Yoo. 2022 · 2022
Earlier work this paper cites.
Factuality challenges in the era of large language models
Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha, Tanmoy Chakraborty, Giovanni Luca Ciampaglia, David Corney, Renee DiResta, Emilio Ferrara, Scott Hale, Alon Y. Halevy, Eduard H. Hovy, Heng Ji, Filippo Menczer, Rubén Míguez, Preslav Nakov, Dietram A. Scheufele, Shivam Sharma, and Giovanni Zagni. 2023 · 2023
Earlier work this paper cites.
Touchstone: Evaluating vision-language models by language models
Shuai Bai, Shusheng Yang, Jinze Bai, Peng Wang, Xing Zhang, Junyang Lin, Xinggang Wang, Chang Zhou, and Jingren Zhou. 2023 · 2023
Earlier work this paper cites.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023 · 2023
Earlier work this paper cites.
Aquallm: Audio question answering data generation using large language models
Swarup Ranjan Behera, Krishna Mohan Injeti, Jaya Sai Kiran Patibandla, Praveen Kumar Pokala, and Balakrishna Reddy Pailla. 2023 · 2023
Cited alongside, same era.
Medbench: A large-scale chinese benchmark for evaluating medical large language models
Yan Cai, Linlin Wang, Ye Wang, Gerard de Melo, Ya Zhang, Yanfeng Wang, and Liang He. 2023 · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung yi Lee. 2023 · 2023
Cited alongside, same era.
Chatlaw: Open-source legal large language model with integrated external knowledge bases
Jiaxi Cui, Zongjia Li, Yang Yan, Bohua Chen, and Li Yuan. 2023 · 2023
Cited alongside, same era.
Emphassess: a prosodic benchmark on assessing emphasis transfer in speech-to-speech models
Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration
Chenyang Lyu, Minghao Wu, Longyue Wang, Xinting Huang, Bingshuai Liu, Zefeng Du, Shuming Shi, and Zhaopeng Tu. 2023 · 2023
Later among the works it cites.
Large language models play starcraft ii: Benchmarks and a chain of summarization approach
Weiyu Ma, Qirui Mi, Xue Yan, Yuqiao Wu, Runji Lin, Haifeng Zhang, and Jun Wang. 2023 · 2023
Later among the works it cites.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. 2023 · 2023
Later among the works it cites.
Eric Melz. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maureen de Seyssel, Antony D’Avirro, Adina Williams, and Emmanuel Dupoux. 2023 · 2023
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023 · 2023
Cited alongside, same era.
Lp-musiccaps: Llm-based pseudo music captioning
SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam. 2023 · 2023
Cited alongside, same era.
How abilities in large language models are affected by supervised fine-tuning data composition
Guanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li, Mingfeng Xue, Dayiheng Liu, Wei Wang, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2023 · 2023
Cited alongside, same era.
Scene graph as pivoting: Inference-time image-free unsupervised multimodal machine translation with visual scene hallucination
Hao Fei, Qianfeng Liu, Meishan Zhang, M. Zhang, and Tat seng Chua. 2023 · 2023
Cited alongside, same era.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Zhen fei Yin, Jiong Wang, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Lu Sheng, Lei Bai, Xiaoshui Huang, Zhiyong Wang, Wanli Ouyang, and Jing Shao. 2023 · 2023
Cited alongside, same era.
Chainpoll: A high efficacy method for llm hallucination detection
Robert Friel and Atindriyo Sanyal. 2023 · 2023
Cited alongside, same era.
Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, and James Glass. 2023 · 2023
Cited alongside, same era.
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Later among the works it cites.
Explanation shift: How did the distribution shift impact the model?
Carlos Mougan, Klaus Broelemann, David Masip, Gjergji Kasneci, Thanassis Thiropanis, and Steffen Staab. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra-Aimée Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Invite: a testbed of automatically generated invalid questions to evaluate large language models for hallucinations
Anil Ramakrishna, Rahul Gupta, Jens Lehmann, and Morteza Ziyadi. 2023 · 2023
Later among the works it cites.
Robots that ask for help: Uncertainty alignment for large language model planners
Allen Z Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, et al. 2023 · 2023
Later among the works it cites.
Delucionqa: Detecting hallucinations in domain-specific question answering
Mobashir Sadat, Zhengyu Zhou, Lukas Lange, Jun Araki, Arsalan Gundroo, Bingqing Wang, Rakesh R Menon, Md. Rizwan Parvez, and Zhe Feng. 2023 · 2023
Later among the works it cites.
Bhaskarjit Sarmah, Tianjie Zhu, Dhagash Mehta, and Stefano Pasquali. 2023 · 2023
Later among the works it cites.
Halp: Hallucinating latent positives for skeleton-based self-supervised learning of actions
Anshul B. Shah, Aniket Basu Roy, Ketul Shah, Shlok Kumar Mishra, David W. Jacobs, Anoop Cherian, and Ramalingam Chellappa. 2023 · 2023
Later among the works it cites.
Entity-augmented code generation
Anton Shapkin, Denis Litvinov, and Timofey Bryksin. 2023 · 2023
Later among the works it cites.
Arvind Krishna Sridhar, Yinyi Guo, Erik Visser, and Rehana Mahfuz. 2023 · 2023
Later among the works it cites.
Behind the magic, merlim: Multi-modal evaluation benchmark for large image-language models
Andrés Villa, Juan Carlos Le’on Alc’azar, Alvaro Soto, and Bernard Ghanem. 2023 · 2023
Later among the works it cites.
Hallucination improves the performance of unsupervised visual representation learning
Jing Wu, Jennifer Hobbs, and Naira Hovakimyan. 2023 · 2023
Later among the works it cites.
Kinematic-aware prompting for generalizable articulated object manipulation with llms
Wenke Xia, Dong Wang, Xincheng Pang, Zhigang Wang, Bin Zhao, and Di Hu. 2023 · 2023
Later among the works it cites.
Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Peng Xu, Wenqi Shao, Kaipeng Zhang, Peng Gao, Shuo Liu, Meng Lei, Fanqing Meng, Siyuan Huang, Yu Jiao Qiao, and Ping Luo. 2023 · 2023
Later among the works it cites.
Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent
Jianing Yang, Xuweiyi Chen, Shengyi Qian, Nikhil Madaan, Madhavan Iyengar, David F Fouhey, and Joyce Chai. 2023 · 2023
Later among the works it cites.
Llm lies: Hallucinations are not bugs, but features as adversarial examples
Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu, Munan Ning, and Li Yuan. 2023 · 2023
Later among the works it cites.
Woodpecker: Hallucination correction for multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, and Enhong Chen. 2023 · 2023
Later among the works it cites.
Halluaudio: Hallucinate frequency as concepts for few-shot audio classification
Zhongjie Yu, Shuyang Wang, Lin Chen, and Zhongwei Cheng. 2023e · 2023
Later among the works it cites.
Alignscore: Evaluating factual consistency with a unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023 · 2023
Later among the works it cites.
User-controlled knowledge fusion in large language models: Balancing creativity and hallucination
Chen Zhang. 2023 · 2023
Later among the works it cites.
Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization
Zhiyuan Zhao, Bin Wang, Linke Ouyang, Xiao wen Dong, Jiaqi Wang, and Conghui He. 2023 · 2023
Later among the works it cites.