Fetching the paper…
Reading the bibliography…
Despite their impressive generative capabilities, LLMs are hindered by fact-conflicting hallucinations in real-world applications.
Wikidata: a free collaborative knowledgebase
Denny Vrandecic and Markus Krötzsch · 2014
Earlier work this paper cites.
Truth, truthiness, triangulation: A news literacy toolkit for a “post-truth” world
Joyce Valenza · 2016
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Handling divergent reference texts when evaluating table-to-text generation
Bhuwan Dhingra, Manaal Faruqui, Ankur P. Parikh, Ming-Wei Chang, Dipanjan Das, and William W. Cohen · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Tabfact: A large-scale dataset for table-based fact verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang · 2020
Earlier work this paper cites.
Climate-fever: A dataset for verification of real-world climate claims
Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold · 2020
Earlier work this paper cites.
Hover: A dataset for many-hop fact extraction and claim verification
Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Kumar Singh, and Mohit Bansal · 2020
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryściński, Bryan McCann, Caiming Xiong, and Richard Socher · 2020
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi · 2020
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi · 2020
Earlier work this paper cites.
FEVEROUS: fact extraction and verification over unstructured and structured information
Rami Aly, Zhijiang Guo, Michael Sejr Schlichtkrull, James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Oana Cocarascu, and Arpit Mittal · 2021
Earlier work this paper cites.
Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov · 2021
Earlier work this paper cites.
Measuring attribution in natural language generation models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter · 2021
Earlier work this paper cites.
Covid-fact: Fact extraction and verification of real-world claims on COVID-19 pandemic
Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan · 2021
Earlier work this paper cites.
Evidence-based fact-checking of health-related claims
Mourad Sarrouti, Asma Ben Abacha, Yassine Mrabet, and Dina Demner-Fushman · 2021
Cited alongside, same era.
Retrieval augmentation reduces hallucination in conversation
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston · 2021
Cited alongside, same era.
Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction
Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen · 2022
Cited alongside, same era.
Faithful reasoning using large language models
Antonia Creswell and Murray Shanahan · 2022
Cited alongside, same era.
Faithdial: A faithful benchmark for information-seeking dialogue
Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar R. Zaïane, Mo Yu, Edoardo Maria Ponti, and Siva Reddy · 2022
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Closest in time.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen · 2023
Closest in time.
Structure guided multi-modal pre-trained transformer for knowledge graph reasoning
Ke Liang, Sihang Zhou, Yue Liu, Lingyuan Meng, Meng Liu, and Xinwang Liu · 2023
Closest in time.
Generating benchmarks for factuality evaluation of language models
Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, and Yoav Shoham · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dialfact: A benchmark for fact-checking in dialogue
Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi · 2022
Cited alongside, same era.
Covert: A corpus of fact-checked biomedical COVID-19 tweets
Isabelle Mohr, Amelie Wührl, and Roman Klinger · 2022
Cited alongside, same era.
Chatgpt: Optimizing language models for dialogue
OpenAI · 2022
Cited alongside, same era.
Building a knowledge graph to enable precision medicine
Payal Chandak, Kexin Huang, and Marinka Zitnik · 2023
Cited alongside, same era.
I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu · 2023
Cited alongside, same era.
Yujia Qin, Shengding Hu, and et al · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Survey on factuality in large language models: Knowledge, retrieval and domain-specificity
Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, and et al · 2023
Closest in time.
Resolving knowledge conflicts in large language models
Yike Wang, Shangbin Feng, Heng Wang, Weijia Shi, Vidhisha Balachandran, Tianxing He, and Yulia Tsvetkov · 2023
Closest in time.
Editing large language models: Problems, methods, and opportunities
Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang · 2023
Closest in time.
Do large language models know what they don’t know?
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang · 2023
Closest in time.
Siren’s song in the AI ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, and et al · 2023
Closest in time.
C-RAG: certified generation risks for retrieval-augmented language models
Mintong Kang, Nezihe Merve Gürel, Ning Yu, Dawn Song, and Bo Li · 2024
Closest in time.
Knowledge graph contrastive learning based on relation-symmetrical structure
Ke Liang, Yue Liu, Sihang Zhou, Wenxuan Tu, Yi Wen, Xihong Yang, Xiangjun Dong, and Xinwang Liu · 2024
Closest in time.