Fetching the paper…
Reading the bibliography…
Large language models (LLMs) exhibit hallucinations in long-form question-answering tasks across various domains and wide applications.
Lin, C.Y.: Rouge: A package for automatic evaluation of summaries. In: Text summarization branches out. pp. 74–81 (2004)
2004
Earlier work this paper cites.
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W.W., Salakhutdinov, R., Manning, C.D.: HotpotQA: A dataset for diverse, explainable multi-hop question answering. In: Conference on Empirical Methods in Natural Language Processing (EMNLP) (2018)
2018
Earlier work this paper cites.
Garg, S., Peitz, S., Nallasamy, U., Paulik, M.: Jointly learning to align and translate with transformer models. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). pp. 4453–4462 (2019)
2019
Earlier work this paper cites.
Durmus, E., He, H., Diab, M.: Feqa: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 5055–5070 (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Weng, R., Yu, H., Wei, X., Luo, W.: Towards enhancing faithfulness for neural machine translation. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). pp. 2675–2684 (2020)
2020
Earlier work this paper cites.
Zhang*, T., Kishore*, V., Wu*, F., Weinberger, K.Q., Artzi, Y.: Bertscore: Evaluating text generation with bert. In: International Conference on Learning Representations (2020), https://openreview.net/forum?id=SkeHuCVFDr
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Mena, J., Pujol, O., Vitrià, J.: A survey on uncertainty estimation in deep learning classification systems from a bayesian perspective. ACM Computing Surveys (CSUR) 54
2021
Earlier work this paper cites.
Petroni, F., Piktus, A., Fan, A., Lewis, P., Yazdani, M., De Cao, N., Thorne, J., Jernite, Y., Karpukhin, V., Maillard, J., et al.: Kilt: a benchmark for knowledge intensive language tasks. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 2523–2544 (2021)
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Dziri, N., Kamalloo, E., Milton, S., Zaiane, O., Yu, M., Ponti, E.M., Reddy, S.: Faithdial: A faithful benchmark for information-seeking dialogue. Transactions of the Association for Computational Linguistics 10
2022
Earlier work this paper cites.
Gupta, P., Wu, C.S., Liu, W., Xiong, C.: Dialfact: A benchmark for fact-checking in dialogue. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 3785–3801 (2022)
2022
Earlier work this paper cites.
He, C., Li, W., Jin, Z., Wang, B., Xu, C., Lin: Opendatalab: Empowering general artificialintelligence with open datasets. https://opendatalab.com (2022)
2022
Earlier work this paper cites.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Madotto, A., Fung, P.: Survey of hallucination in natural language generation. ACM Computing Surveys (2022)
2022
Earlier work this paper cites.
Laban, P., Schnabel, T., Bennett, P.N., Hearst, M.A.: Summac: Re-visiting nli-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics 10
2022
Earlier work this paper cites.
Rebuffel, C., Roberti, M., Soulier, L., Scoutheeten, G., Cancelliere, R., Gallinari, P.: Controlling hallucinations at word level in data-to-text generation. Data Mining and Knowledge Discovery 36
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Semnani, S., Yao, V., Zhang, H., Lam, M.: Wikichat: Stopping the hallucination of large language model chatbots by few-shot grounding on wikipedia. In: Findings of the Association for Computational Linguistics: EMNLP 2023. pp. 2387–2413 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., Xin, R.: Free dolly: Introducing the world’s first truly open instruction-tuned llm (2023), https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm
2023
Cited alongside, same era.
Contributors, L.: Lmdeploy: A toolkit for compressing, deploying, and serving llm. https://github.com/InternLM/lmdeploy (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Gilardi, F., Alizadeh, M., Kubli, M.: Chatgpt outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences 120
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Ji, Z., Liu, Z., Lee, N., Yu, T., Wilie, B., Zeng, M., Fung, P.: RHO: Reducing hallucination in open-domain dialogues with knowledge grounding. In: Rogers, A., Boyd-Graber, J., Okazaki, N. (eds.) Findings of the Association for Computational Linguistics: ACL 2023. pp. 4504–4522. Association for Computational Linguistics, Toronto, Canada (Jul 2023). https://doi.org/10.18653/v1/2023.findings-acl.275, https://aclanthology.org/2023.findings-acl.275
2023
Cited alongside, same era.
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
He, Z., Huang, C.Y., Ding, C.K.C., Rohatgi, S., Huang, T.H.K.: If in a crowdsourced data annotation pipeline, a gpt-4. In: Proceedings of the CHI Conference on Human Factors in Computing Systems. pp. 1–25 (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., Gan, C.: Principle-driven self-alignment of language models from scratch with minimal human supervision. Advances in Neural Information Processing Systems 36
2024
Closest in time.
2024
Closest in time.