Fetching the paper…
Reading the bibliography…
While hallucinations of large language models (LLMs) prevail as a major challenge, existing evaluation benchmarks on factuality do not cover the diverse domains of knowledge that the real-world users of LLMs seek information about.
NLTK: The natural language toolkit
Steven Bird and Edward Loper · 2004
Earlier work this paper cites.
Disinformation in the online information ecosystem: Detection, mitigation and challenges, 2020
Amrita Bhattacharjee, Kai Shu, Min Gao, and Huan Liu · 2020
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab · 2020
Earlier work this paper cites.
Evaluating factuality in generation with dependency-level entailment
Tanya Goyal and Greg Durrett · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald · 2020
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis · 2020
Earlier work this paper cites.
The rise and fall of fake news sites: A traffic analysis, 2021
Manolis Chalkiadakis, Alexandros Kornilakis, Panagiotis Papadopoulos, Evangelos P. Markatos, and Nicolas Kourtellis · 2021
Earlier work this paper cites.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong · 2022
Earlier work this paper cites.
Summac: Re-visiting nli-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N Bennett, and Marti A Hearst · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Enabling large language models to generate text with citations
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen · 2023
Earlier work this paper cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
HaluEval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen · 2023
Cited alongside, same era.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation, 2023
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
The shifted and the overlooked: A task-oriented investigation of user-GPT interactions
Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Evaluating llms at detecting errors in llm responses
Ryo Kamoi, Sarkar Snigdha Sarathi Das, Renze Lou, Jihyun Janice Ahn, Yilun Zhao, Xiaoxin Lu, Nan Zhang, Yusen Zhang, Ranran Haoran Zhang, Sujeeth Reddy Vummanthala, et al · 2024
Closest in time.
Expertqa: Expert-curated questions and attributed answers, 2024
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth · 2024
Closest in time.
Fine-grained hallucination detection and editing for language models, 2024
Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang, Graham Neubig, Yulia Tsvetkov, and Hannaneh Hajishirzi · 2024
Closest in time.
Fine-grained hallucination detection and editing for language models
Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang, Graham Neubig, Yulia Tsvetkov, and Hannaneh Hajishirzi · 2024
Closest in time.
Entity Types
Wolfram Language & System · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Command r+ model card
Cohere For AI · 2024
Cited alongside, same era.
Llama 3 model card
AI@Meta · 2024
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Cited alongside, same era.
Large language models and user trust: Consequence of self-referential learning loop and the deskilling of health care professionals
Avishek Choudhury and Zaira Chaudhry · 2024
Cited alongside, same era.
The hallucinations leaderboard – an open effort to measure hallucinations in large language models, 2024
Giwon Hong, Aryo Pradipta Gema, Rohit Saxena, Xiaotang Du, Ping Nie, Yu Zhao, Laura Perez-Beltrachini, Max Ryabinin, Xuanli He, Clémentine Fourrier, and Pasquale Minervini · 2024
Cited alongside, same era.
Tofueval: Evaluating hallucinations of llms on topic-focused dialogue summarization, 2024
Liyan Tang, Igor Shalyminov, Amy Wing mei Wong, Jon Burnsky, Jake W. Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, and Kathleen McKeown · 2024
Closest in time.
Worse than random? an embarrassingly simple probing evaluation of large multimodal models in medical vqa, 2024
Qianqi Yan, Xuehai He, Xiang Yue, and Xin Eric Wang · 2024
Closest in time.
Wildchat: 1m chatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng · 2024
Closest in time.
Felm: Benchmarking factuality evaluation of large language models
Yiran Zhao, Jinghan Zhang, I Chern, Siyang Gao, Pengfei Liu, Junxian He, et al · 2024
Closest in time.
Halueval-wild: Evaluating hallucinations of language models in the wild
Zhiying Zhu, Zhiqing Sun, and Yiming Yang · 2024
Closest in time.