Fetching the paper…
Reading the bibliography…
Evaluating the factuality of long-form large language model (LLM)-generated text is an important challenge.
Reading wikipedia to answer open-domain questions
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes · 2017
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Reasoning over semantic-level graph for fact checking
Wanjun Zhong, Jingjing Xu, Duyu Tang, Zenan Xu, Nan Duan, Ming Zhou, Jiahai Wang, and Jian Yin · 2019
Earlier work this paper cites.
Kilt: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, et al · 2020
Earlier work this paper cites.
X-fact: A new benchmark dataset for multilingual fact checking
Ashim Gupta and Vivek Srikumar · 2021
Earlier work this paper cites.
The importance of modeling social factors of language: Theory and practice
Dirk Hovy and Diyi Yang · 2021
Earlier work this paper cites.
mface: Multilingual summarization with factual consistency evaluation
Roee Aharoni, Shashi Narayan, Joshua Maynez, Jonathan Herzig, Elizabeth Clark, and Mirella Lapata · 2022
Earlier work this paper cites.
A survey on automated fact-checking
Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos · 2022
Earlier work this paper cites.
Challenges and strategies in cross-cultural NLP
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard · 2022
Earlier work this paper cites.
Proofver: Natural logic theorem proving for fact verification
Amrith Krishna, Sebastian Riedel, and Andreas Vlachos · 2022
Earlier work this paper cites.
Nonparametric masked language modeling
Sewon Min, Weijia Shi, Mike Lewis, Xilun Chen, Wen-tau Yih, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
MEGA: Multilingual evaluation of generative AI
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram · 2023
Cited alongside, same era.
The internal state of an llm knows when it’s lying, 2023
Amos Azaria and Tom Mitchell · 2023
Cited alongside, same era.
A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung · 2023
Cited alongside, same era.
Multilingual large language models leak human stereotypes across language boundaries, 2023
Yang Trista Cao, Anna Sotnikova, Jieyu Zhao, Linda X. Zou, Rachel Rudinger, and Hal Daume III au2 · 2023
Cited alongside, same era.
I Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, Pengfei Liu, et al · 2023
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models, 2023
Potsawee Manakul, Adian Liusie, and Mark J. F. Gales · 2023
Later among the works it cites.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation, 2023
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Having beer after prayer? measuring cultural bias in large language models
Tarek Naous, Michael J Ryan, and Wei Xu · 2023
Later among the works it cites.
Factcheck-gpt: End-to-end fine-grained document-level fact-checking and correction of llm output
Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid, Arnav Arora, Aleksandr Rubashevskii, Jiahui Geng, Osama Mohammed Afzal, Liangming Pan, Nadav Borenstein, Aditya Pillai, et al · 2023
Later among the works it cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models, 2023
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston · 2023
Cited alongside, same era.
Culturally aware natural language inference
Jing Huang and Diyi Yang · 2023
Cited alongside, same era.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Is chatgpt a good translator? a preliminary study
Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu · 2023
Cited alongside, same era.
Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries, 2023
Yiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu, Munmun De Choudhury, and Srijan Kumar · 2023
Cited alongside, same era.
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2023
Later among the works it cites.
Comparing hallucination detection metrics for multilingual generation
Haoqiang Kang, Terra Blevins, and Luke Zettlemoyer · 2024
Closest in time.
Global-liar: Factuality of llms over time and geographic regions
Shujaat Mirza, Bruno Coelho, Yuyuan Cui, Christina Pöpper, and Damon McCoy · 2024
Closest in time.
Fine-grained hallucination detection and editing for language models
Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang, Graham Neubig, Yulia Tsvetkov, and Hannaneh Hajishirzi · 2024
Closest in time.
Long-form factuality in large language models, 2024
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, and Quoc V. Le · 2024
Closest in time.
Wikipedia:multilingual statistics — Wikipedia, the free encyclopedia, 2024
Wikipedia contributors · 2024
Closest in time.