Fetching the paper…
Reading the bibliography…
In many real-world scenarios, a single Large Language Model (LLM) may encounter contradictory claims-some accurate, others forcefully incorrect-and must judge which is true.
G. Irving, P. F. Christiano, and D. Amodei · 2018
Earlier work this paper cites.
How can we know when language models know? On the calibration of language models for QA
Z. Jiang, P. Xia, D. Misra, Q. Lei, N. Shalom, R. Nallapati, et al · 2021
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
Language models (mostly) know what they know
S. Kadavath, T. Conerly, A. Askell, et al · 2022
Earlier work this paper cites.
Debate helps supervise unreliable experts
J. Michael, S. Mahdi, D. Rein, et al · 2023
Cited alongside, same era.
ChatEval: Towards better LLM-based evaluators through multi-agent debate
C. M. Chan, W. Z. Chen, Y. S. Su, and X. Ma · 2023
Cited alongside, same era.
GPT-4 technical report
OpenAI · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Z. Ji, et al · 2023
Cited alongside, same era.
On scalable oversight with weak LLMs judging strong LLMs
Z. Kenton, N. Y. Siegel, J. Kramar, et al · 2024
Later among the works it cites.
Adversarial multi-agent evaluation of large language models through iterative debates
C. Bandi and A. Harrasse · 2024
Later among the works it cites.
"The Earth is Flat because…": Investigating LLMs’ belief towards misinformation via persuasive conversation
W. L. Chiang, Z. Li, Z. Lin, Y. Sheng, et al · 2024
Later among the works it cites.
The persuasive power of large language models
S. M. Breum, D. V. Egdal, V. G. Mortensen, A. G. Moller, and L. M. Aiello · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…