Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have revolutionized Natural Language Processing (NLP) based applications including automated text generation, question answering, chatbots, and others.
Bertscore: Evaluating text generation with bert,
T. Zhang*, V. Kishore*, F. Wu*, K. Q. Weinberger, Y. Artzi, · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks,
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al., · 2020
Earlier work this paper cites.
Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation,
Y. Sun, S. Wang, S. Feng, S. Ding, C. Pang, J. Shang, J. Liu, X. Chen, Y. Zhao, Y. Lu, et al., · 2021
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation,
W. Yuan, G. Neubig, P. Liu, · 2021
Earlier work this paper cites.
Adapters for enhanced modeling of multilingual knowledge and text,
Y. Hou, W. Jiao, M. Liu, C. Allen, Z. Tu, M. Sachan, · 2022
Earlier work this paper cites.
FactGraph: Evaluating factuality in summarization with semantic graph representations,
L. F. R. Ribeiro, M. Liu, I. Gurevych, M. Dreyer, M. Bansal, · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods,
S. Lin, J. Hilton, O. Evans, · 2022
Earlier work this paper cites.
L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al., · 2023
Earlier work this paper cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models,
Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen, et al., · 2023
Earlier work this paper cites.
Large Language Models and Knowledge Graphs: Opportunities and Challenges,
J. Z. Pan, S. Razniewski, J.-C. Kalo, S. Singhania, J. Chen, S. Dietze, H. Jabeen, J. Omeliyanenko, W. Zhang, M. Lissandrini, R. Biswas, G. de Melo, A. Bonifati, E. Vakaj, M. Dragoni, D. Graux, · 2023
Earlier work this paper cites.
Knowledge graph enhanced language models for sentiment analysis,
J. Li, X. Li, L. Hu, Y. Zhang, J. Wang, · 2023
Earlier work this paper cites.
Dola: Decoding by contrasting layers improves factuality in large language models,
Y.-S. Chuang, Y. Xie, H. Luo, Y. Kim, J. Glass, P. He, · 2023
Earlier work this paper cites.
Think-on-graph: Deep and responsible reasoning of large language model with knowledge graph,
J. Sun, C. Xu, L. Tang, S. Wang, C. Lin, Y. Gong, H.-Y. Shum, J. Guo, · 2023
Earlier work this paper cites.
Knowledge injection to counter large language model (llm) hallucination,
A. Martino, M. Iannelli, C. Truong, · 2023
Earlier work this paper cites.
Med-HALT: Medical domain hallucination test for large language models,
A. Pal, L. K. Umapathi, M. Sankarasubbu, · 2023
Earlier work this paper cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models,
J. Li, X. Cheng, X. Zhao, J.-Y. Nie, J.-R. Wen, · 2023
Earlier work this paper cites.
FLEEK: Factual error detection and correction with evidence retrieved from external knowledge,
F. Fatahi Bayat, K. Qian, B. Han, Y. Sang, A. Belyy, S. Khorshidi, F. Wu, I. Ilyas, Y. Li, · 2023
Cited alongside, same era.
Enhancing uncertainty-based hallucination detection with stronger focus,
T. Zhang, L. Qiu, Q. Guo, C. Deng, Y. Zhang, Z. Zhang, C. Zhou, X. Wang, L. Fu, · 2023
Cited alongside, same era.
Cross-lingual consistency of factual knowledge in multilingual language models,
J. Qi, R. Fernández, A. Bisazza, · 2023
Cited alongside, same era.
Multilingual Knowledge Graphs and Low-Resource Languages: A Review,
L.-A. Kaffee, R. Biswas, C. M. Keet, E. K. Vakaj, G. de Melo, · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al., · 2023
Cited alongside, same era.
State of what art? a call for multi-prompt llm evaluation,
M. Mizrahi, G. Kaplan, D. Malkin, R. Dror, D. Shahaf, G. Stanovsky, · 2024
Closest in time.
Defan: Definitive answer dataset for llms hallucination evaluation,
A. Rahman, S. Anwar, M. Usman, A. Mian, · 2024
Closest in time.
SemEval-2024 task 6: SHROOM, a shared-task on hallucinations and related observable overgeneration mistakes,
T. Mickus, E. Zosa, R. Vazquez, T. Vahtola, J. Tiedemann, V. Segonne, A. Raganato, M. Apidianaki, · 2024
Closest in time.
2024
Closest in time.
OpenAI, Measuring short-form factuality in large language models, 2024. URL: https://cdn.openai.com/papers/simpleqa.pdf
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. Koh, M. Iyyer, L. Zettlemoyer, H. Hajishirzi, · 2023
Cited alongside, same era.
Factuality challenges in the era of large language models and opportunities for fact-checking,
I. Augenstein, T. Baldwin, M. Cha, T. Chakraborty, G. L. Ciampaglia, D. Corney, R. DiResta, E. Ferrara, S. Hale, A. Halevy, et al., · 2024
Cited alongside, same era.
AI ‘news’ content farms are easy to make and hard to detect: A case study in Italian,
G. Puccetti, A. Rogers, C. Alzetta, F. Dell’Orletta, A. Esuli, · 2024
Cited alongside, same era.
Hallucinations in llms: Understanding and addressing challenges,
G. Perković, A. Drobnjak, I. Botički, · 2024
Cited alongside, same era.
Unifying large language models and knowledge graphs: A roadmap,
S. Pan, L. Luo, Y. Wang, C. Chen, J. Wang, X. Wu, · 2024
Cited alongside, same era.
Trusting your evidence: Hallucinate less with context-aware decoding,
W. Shi, X. Han, M. Lewis, Y. Tsvetkov, L. Zettlemoyer, W.-t. Yih, · 2024
Cited alongside, same era.
KG-adapter: Enabling knowledge graph integration in large language models through parameter-efficient fine-tuning,
S. Tian, Y. Luo, T. Xu, C. Yuan, H. Jiang, C. Wei, X. Wang, · 2024
Cited alongside, same era.
Closest in time.
Hallucination is inevitable: An innate limitation of large language models,
Z. Xu, S. Jain, M. Kankanhalli, · 2024
Closest in time.
Llms will always hallucinate, and we need to live with this,
S. Banerjee, A. Agarwal, S. Singla, · 2024
Closest in time.
FactAlign: Fact-level hallucination detection and classification through knowledge graph alignment,
M. Rashad, A. Zahran, A. Amin, A. Abdelaal, M. Altantawy, · 2024
Closest in time.
Query in your tongue: Reinforce large language models with retrievers for cross-lingual search generative experience,
P. Guo, Y. Hu, Y. Cao, Y. Ren, Y. Li, H. Huang, · 2024
Closest in time.
Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries,
Y. Jin, M. Chandra, G. Verma, Y. Hu, M. De Choudhury, S. Kumar, · 2024
Closest in time.
Unifying local and global knowledge: Empowering large language models as political experts with knowledge graphs,
X. Mou, Z. Li, H. Lyu, J. Luo, Z. Wei, · 2024
Closest in time.
Generative expression constrained knowledge-based decoding for open data,
L. Lageweg, B. Kruit, · 2024
Closest in time.
Multilingual hallucination gaps in large language models,
C. Chataigner, A. Taïk, G. Farnadi, · 2024
Closest in time.
2024
Closest in time.
R. Vázquez, T. Mickus, E. Zosa, T. Vahtola, J. Tiedemann, A. Sinha, V. Segonne, F. Sánchez-Vega, A. Raganato, J. Karlgren, S. Ji, L. Guillou, J. Attieh, M. Apidianaki, SemEval-2025 Task 3: Mu-SHROOM, the multilingual shared-task on hallucinations and related observable overgeneration mistakes, 2025. URL: https://helsinki-nlp.github.io/shroom/
2025
Closest in time.