Fetching the paper…
Reading the bibliography…
Hallucinations in LLMs pose a significant concern to their safe deployment in real-world applications.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Directional statistics
Mardia, K. V. and Jupp, P. E · 2009
Earlier work this paper cites.
Optimal transport: old and new
Villani, C. et al · 2009
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Cuturi, M · 2013
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Object hallucination in image captioning
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K · 2018
Earlier work this paper cites.
Flexible imputation of missing data
Van Buuren, S · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A · 2020
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A. P · 2020
Earlier work this paper cites.
Uncertainty estimation in autoregressive structured prediction
Malinin, A. and Gales, M · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al · 2022
Earlier work this paper cites.
Out-of-distribution detection and selective generation for conditional language models
Ren, J., Luo, J., Zhao, Y., Krishna, K., Saleh, M., Lakshminarayanan, B., and Liu, P. J · 2022
Earlier work this paper cites.
Pico: Contrastive label disambiguation for partial label learning
Wang, H., Xiao, R., Li, Y., Feng, L., Niu, G., Chen, G., and Zhao, J · 2022
Cited alongside, same era.
The internal state of an llm knows when it’s lying
Azaria, A. and Mitchell, T · 2023
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2023
Cited alongside, same era.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., et al · 2023
Cited alongside, same era.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Kuhn, L., Gal, Y., and Farquhar, S · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Later among the works it cites.
Semantic entropy probes: Robust and cheap hallucination detection in llms
Kossen, J., Han, J., Razzak, M., Schut, L., Malik, S., and Gal, Y · 2024
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M · 2024
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Lin, Z., Trivedi, S., and Sun, J · 2024
Later among the works it cites.
Mitigating hallucination in large multi-modal models via robust instruction tuning
Liu, F., Lin, K., Li, L., Wang, J., Yacoob, Y., and Wang, L · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Manakul, P., Liusie, A., and Gales, M. J · 2023
Cited alongside, same era.
Med-halt: Medical domain hallucination test for large language models
Pal, A., Umapathi, L. K., and Sankarasubbu, M · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., et al · 2023
Cited alongside, same era.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Cited alongside, same era.
Later among the works it cites.
The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Marks, S. and Tegmark, M · 2024
Later among the works it cites.
Prism: A framework for decoupling and assessing the capabilities of vlms
Qiao, Y., Duan, H., Fang, X., Yang, J., Chen, L., Zhang, S., Wang, J., Lin, D., and Chen, K · 2024
Later among the works it cites.
Aligning large multimodal models with factually augmented RLHF
Sun, Z., Shen, S., Cao, S., Liu, H., Li, C., Shen, Y., Gan, C., Gui, L., Wang, Y.-X., Yang, Y., Keutzer, K., and Darrell, T · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Tong, P., Brown, E., Wu, P., Woo, S., IYER, A. J. V., Akula, S. C., Yang, S., Yang, J., Middepogu, M., Wang, Z., et al · 2024
Later among the works it cites.
Reft: Representation finetuning for language models
Wu, Z., Arora, A., Wang, Z., Geiger, A., Jurafsky, D., Manning, C. D., and Potts, C · 2024
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Ict: Image-object cross-level trusted intervention for mitigating object hallucination in large vision-language models
Chen, J., Zhang, T., Huang, S., Niu, Y., Zhang, L., Wen, L., and Hu, X · 2025
Closest in time.
Truthprint: Mitigating lvlm object hallucination via latent truthful-guided pre-intervention
Duan, J., Kong, F., Cheng, H., Diffenderfer, J., Kailkhura, B., Sun, L., Zhu, X., Shi, X., and Xu, K · 2025
Closest in time.
A unified understanding and evaluation of steering methods
Im, S. and Li, Y · 2025
Closest in time.
Reducing hallucinations in vision-language models via latent space steering
Liu, S., Ye, H., Xing, L., and Zou, J · 2025
Closest in time.
Nullu: Mitigating object hallucinations in large vision-language models via halluspace projection
Yang, L., Zheng, Z., Chen, B., Zhao, Z., Lin, C., and Shen, C · 2025
Closest in time.
Can your uncertainty scores detect hallucinated entity?
Yeh, M.-H., Kamachee, M., Park, S., and Li, Y · 2025
Closest in time.