Fetching the paper…
Reading the bibliography…
We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Earlier work this paper cites.
Language models (mostly) know what they know, 2022
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, S. Johnston, S. El-Showk, A. Jones, N. Elhage, T. Hume, A. Chen, Y. Bai, S. Bowman, S. Fort, D. Ganguli, D. Hernandez, J. Jacobson, J. Kernion, S. Kravec, L. Lovitt, K. Ndousse, C. Olsson, S. Ringer, D. Amodei, T. Brown, J. Clark, N. Joseph, B. Mann, S. McCandlish, C. Olah, and J. Kaplan · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
Evaluating hallucinations in chinese large language models
Q. Cheng, T. Sun, W. Zhang, S. Wang, X. Liu, M. Zhang, J. He, M. Huang, Z. Yin, K. Chen, et al · 2023
Earlier work this paper cites.
HaluEval: A large-scale hallucination evaluation benchmark for large language models
J. Li, X. Cheng, X. Zhao, J.-Y. Nie, and J.-R. Wen · 2023
Cited alongside, same era.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi · 2023
Cited alongside, same era.
Freshllms: Refreshing large language models with search engine augmentation, 2023
T. Vu, M. Iyyer, X. Wang, N. Constant, J. Wei, J. Wei, C. Tar, Y.-H. Sung, D. Zhou, Q. Le, and T. Luong · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Cited alongside, same era.
Do language models know when they’re hallucinating references?
Claude 3 model card, 2024
P. Anthropic · 2024
Closest in time.
Personal communication, July 2024
A. T. Kalai · 2024
Closest in time.
Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation
S. Krishna, K. Krishna, A. Mohananey, S. Schwarcz, A. Stambler, S. Upadhyay, and M. Faruqui · 2024
Closest in time.
Long-form factuality in large language models
J. Wei, C. Yang, X. Song, Y. Lu, N. Hu, J. Huang, D. Tran, D. Peng, R. Liu, D. Huang, C. Du, and Q. V. Le · 2024
Closest in time.
FELM: Benchmarking factuality evaluation of large language models
Y. Zhao, J. Zhang, I. Chern, S. Gao, P. Liu, J. He, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Agrawal, M. Suzgun, L. Mackey, and A. T. Kalai · 2024
Cited alongside, same era.
Teaching models to express their uncertainty in words
S. Lin, J. Hilton, and O. Evans
Cited in the paper.
Hello gpt-4o, 2024a
OpenAI
Cited in the paper.
Openai o1-mini, 2024b
OpenAI
Cited in the paper.
Learning to reason with llms, 2024c
OpenAI
Cited in the paper.