Fetching the paper…
Reading the bibliography…
The increased use of large language models (LLMs) across a variety of real-world applications calls for automatic tools to check the factual accuracy of their outputs, as LLMs often hallucinate.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021 · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Attributed text generation via post-hoc research and revision
Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al. 2022 · 2022
Earlier work this paper cites.
Factuality enhanced language models for open-ended text generation
Nayeon Lee, Wei Ping, and Peng et al. Xu. 2022 · 2022
Earlier work this paper cites.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023 · 2023
Earlier work this paper cites.
A categorical archive of ChatGPT failures
Ali Borji. 2023 · 2023
Earlier work this paper cites.
FELM: Benchmarking factuality evaluation of large language models
Shiqi Chen, Yiran Zhao, Jinghan Zhang, I-Chun Chern, Siyang Gao, Pengfei Liu, and Junxian He. 2023 · 2023
Cited alongside, same era.
I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
DoLa: Decoding by contrasting layers improves factuality in large language models
Yung-Sung Chuang, Yujia Xie, and Hongyin Luo et al. 2023 · 2023
Cited alongside, same era.
A survey of language model confidence estimation and calibration
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2023 · 2023
Cited alongside, same era.
Trusting your evidence: Hallucinate less with context-aware decoding
Weijia Shi, Xiaochuang Han, and et al. 2023 · 2023
Later among the works it cites.
FreshLLMs: Refreshing large language models with search engine augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, et al. 2023 · 2023
Later among the works it cites.
Factcheck-GPT: End-to-end fine-grained document-level fact-checking and correction of llm output
Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid, Arnav Arora, Aleksandr Rubashevskii, Jiahui Geng, Osama Mohammed Afzal, Liangming Pan, Nadav Borenstein, Aditya Pillai, et al. 2023 · 2023
Later among the works it cites.
Do large language models know what they don’t know?
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Llm failure archive (chatgpt and beyond)
Guiven. 2023 · 2023
Cited alongside, same era.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2023
Cited alongside, same era.
Loki: An open-source tool for fact verification
Hao Wang, Yuxia Wang, Minghan Wang, Yilin Geng, Zhen Zhao, Zenan Zhai, Preslav Nakov, Timothy Baldwin, Xudong Han, and Haonan Li. 2024a
Cited in the paper.
OpenFactCheck: A unified framework for factuality evaluation of llms
Yuxia Wang, Minghan Wang, Hasan Iqbal, Georgi Georgiev, Jiahui Geng, and Preslav Nakov. 2024b
Cited in the paper.
How language model hallucinations can snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith. 2023a
Cited in the paper.
Siren’s song in the AI ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023b
Cited in the paper.
Yuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Georgi Georgiev, Rocktim Jyoti Das, and Preslav Nakov. 2024c · 2024
Closest in time.
Long-form factuality in large language models
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, and Quoc V. Le. 2024 · 2024
Closest in time.
FIRE: Fact-checking with iterative retrieval and verification
Zhuohan Xie, Rui Xing, Yuxia Wang, Jiahui Geng, Hasan Iqbal, Dhruv Sahnan, Iryna Gurevych, and Preslav Nakov. 2024 · 2024
Closest in time.