Fetching the paper…
Reading the bibliography…
We introduce FACTS Grounding, an online leaderboard and associated benchmark that evaluates language models' ability to generate text that is factually accurate with respect to given context in the user prompt.
TRUE: Re-evaluating factual consistency evaluation
O. Honovich, R. Aharoni, J. Herzig, H. Taitelbaum, D. Kukliansy, V. Cohen, T. Scialom, I. Szpektor, A. Hassidim, and Y. Matias · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
LongDocFACTScore: Evaluating the factuality of long document abstractive summarisation
J. A. Bishop, Q. Xie, and S. Ananiadou · 2023
Earlier work this paper cites.
Booookscore: A systematic exploration of book-length summarization in the era of LLMs
Y. Chang, K. Lo, T. Goyal, and M. Iyyer · 2023
Earlier work this paper cites.
Trueteacher: Learning factual consistency evaluation with large language models
Z. Gekhman, J. Herzig, R. Aharoni, C. Elkind, and I. Szpektor · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team: R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al · 2023
Earlier work this paper cites.
LongEval: Guidelines for human evaluation of faithfulness in long-form summarization
K. Krishna, E. Bransom, B. Kuehl, M. Iyyer, P. Dasigi, A. Cohan, and K. Lo · 2023
Earlier work this paper cites.
Factgen: Faithful text generation by factuality-aware pre-training and contrastive ranking fine-tuning
Z. Lan, W. Li, J. Su, X. Xiao, J. Liu, W. Wu, and Y. Lyu · 2023
Earlier work this paper cites.
HaluEval: A large-scale hallucination evaluation benchmark for large language models
J. Li, X. Cheng, W. X. Zhao, J.-Y. Nie, and J.-R. Wen · 2023
Earlier work this paper cites.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi · 2023
Earlier work this paper cites.
Fact-checking complex claims with program-guided reasoning
L. Pan, X. Wu, X. Lu, A. T. Luu, W. Y. Wang, M.-Y. Kan, and P. Nakov · 2023
Earlier work this paper cites.
Measuring attribution in natural language generation models
H. Rashkin, V. Nikolaev, M. Lamm, L. Aroyo, M. Collins, D. Das, S. Petrov, G. S. Tomar, I. Turc, and D. Reitter · 2023
Earlier work this paper cites.
Factually consistent summarization via reinforcement learning with textual entailment feedback
P. Roit, J. Ferret, L. Shani, R. Aharoni, G. Cideron, R. Dadashi, M. Geist, S. Girgin, L. Hussenot, O. Keller, et al · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Cited alongside, same era.
The Claude 3 model family: Opus, Sonnet, Haiku, 2024
Anthropic · 2024
Cited alongside, same era.
Introducing gemini 2.0: our new ai model for the agentic era, 2024
Gemini Team · 2024
Cited alongside, same era.
FactAlign: Long-form factuality alignment of large language models
C.-W. Huang and Y.-N. Chen · 2024
Cited alongside, same era.
Do automatic factuality metrics measure factuality? A critical evaluation
S. Ramprasad and B. C. Wallace · 2024
Later among the works it cites.
Evaluating the factuality of zero-shot summarizers across varied domains
S. Ramprasad, K. Krishna, Z. C. Lipton, and B. C. Wallace · 2024
Later among the works it cites.
Data contamination report from the 2024 CONDA shared task
O. Sainz, I. García-Ferrero, A. Jacovi, J. A. Campos, Y. Elazar, E. Agirre, Y. Goldberg, W.-L. Chen, J. Chim, L. Choshen, et al · 2024
Later among the works it cites.
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
Y. Song, Y. Kim, and M. Iyyer · 2024
Later among the works it cites.
Unsupervised real-time hallucination detection based on the internal states of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Jacovi, M. Ambar, E. Ben-David, U. Shaham, A. Feder, M. Geva, D. Marcus, and A. Caciularu · 2024
Cited alongside, same era.
Calibrated language models must hallucinate
A. T. Kalai and S. S. Vempala · 2024
Cited alongside, same era.
One thousand and one pairs: A "novel" challenge for long-context language models
M. Karpinska, K. Thai, K. Lo, T. Goyal, and M. Iyyer · 2024
Cited alongside, same era.
FABLES: Evaluating faithfulness and content selection in book-length summarization
Y. Kim, Y. Chang, M. Karpinska, A. Garimella, V. Manjunatha, K. Lo, T. Goyal, and M. Iyyer · 2024
Cited alongside, same era.
LLM hallucination reasoning with zero-shot knowledge test
S. Lee, H. Hsu, and C.-F. Chen · 2024
Cited alongside, same era.
LLMs as narcissistic evaluators: When ego inflates evaluation scores
Y. Liu, N. Moosavi, and C. Lin · 2024
Cited alongside, same era.
Learning to reason with LLMs, 2024
OpenAI · 2024
Cited alongside, same era.
W. Su, C. Wang, Q. Ai, Y. Hu, Z. Wu, Y. Zhou, and Y. Liu · 2024
Later among the works it cites.
MiniCheck: Efficient fact-checking of LLMs on grounding documents
L. Tang, P. Laban, and G. Durrett · 2024
Later among the works it cites.
Hallucination evaluation model (revision 7437011), 2024
Vectara · 2024
Later among the works it cites.
Self-preference bias in LLM-as-a-judge
K. Wataoka, T. Takahashi, and R. Ri · 2024
Later among the works it cites.
Pride and prejudice: LLM amplifies self-bias in self-refinement
W. Xu, G. Zhu, X. Zhao, L. Pan, L. Li, and W. Wang · 2024
Later among the works it cites.
Justice or prejudice? quantifying biases in llm-as-a-judge
J. Ye, Y. Wang, Y. Huang, D. Chen, Q. Zhang, N. Moniz, T. Gao, W. Geyer, C. Huang, P.-Y. Chen, N. V. Chawla, and X. Zhang · 2024
Later among the works it cites.
HaluEval-Wild: Evaluating hallucinations of language models in the wild
Z. Zhu, Y. Yang, and Z. Sun · 2024
Later among the works it cites.