Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable capabilities, revolutionizing the integration of AI in daily life applications.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., Evans, O.: · 2021
Earlier work this paper cites.
A comprehensive overview of large language models
Naveed, H., Khan, A.U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Barnes, N., Mian, A.: · 2023
Earlier work this paper cites.
A survey of hallucination in large foundation models
Rawte, V., Sheth, A., Das, A.: · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A., Fung, P.: · 2023
Earlier work this paper cites.
Evaluating correctness and faithfulness of instruction-following models for question answering
Adlakha, V., BehnamGhader, P., Lu, X.H., Meade, N., Reddy, S.: · 2023
Earlier work this paper cites.
Generating benchmarks for factuality evaluation of language models
Muhlgay, D., Ram, O., Magar, I., Levine, Y., Ratner, N., Belinkov, Y., Abend, O., Leyton-Brown, K., Shashua, A., Shoham, Y.: · 2023
Earlier work this paper cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., et al.: · 2023
Earlier work this paper cites.
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., et al.: · 2023
Earlier work this paper cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Li, J., Cheng, X., Zhao, W.X., Nie, J.Y., Wen, J.R.: · 2023
Cited alongside, same era.
A new benchmark and reverse validation method for passage-level hallucination detection
Yang, S., Sun, R., Wan, X.: · 2023
Cited alongside, same era.
Zhang, J., Li, Z., Das, K., Malin, B.A., Kumar, S.: · 2023
Cited alongside, same era.
Evaluating hallucinations in chinese large language models
Cheng, Q., Sun, T., Zhang, W., Wang, S., Liu, X., Zhang, M., He, J., Huang, M., Yin, Z., Chen, K., et al.: · 2023
Cited alongside, same era.
The alignment handbook
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Huang, S., Rasul, K., Rush, A.M., Wolf, T.: · 2023
Halueval-wild: Evaluating hallucinations of language models in the wild (2024)
Zhu, Z., Yang, Y., Sun, Z.: · 2024
Closest in time.
Realtime qa: What’s the answer right now?
Kasai, J., Sakaguchi, K., Le Bras, R., Asai, A., Yu, X., Radev, D., Smith, N.A., Choi, Y., Inui, K., et al.: · 2024
Closest in time.
Li, N., Li, Y., Liu, Y., Shi, L., Wang, K., Wang, H.: · 2024
Closest in time.
The dawn after the dark: An empirical study on factuality hallucination in large language models
Li, J., Chen, J., Ren, R., Cheng, X., Zhao, W.X., Nie, J.Y., Wen, J.R.: · 2024
Closest in time.
Jiang, A.Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D.S., Casas, D.d.l., Hanna, E.B., Bressand, F., et al.: · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models (2023)
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al.: · 2023
Cited alongside, same era.
Towards understanding sycophancy in language models
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S.R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S.R., et al.: · 2023
Cited alongside, same era.
Felm: Benchmarking factuality evaluation of large language models
Zhao, Y., Zhang, J., Chern, I., Gao, S., Liu, P., He, J., et al.: · 2024
Cited alongside, same era.
Closest in time.
Meta llama 3
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al.: · 2024
Closest in time.
Google ai for developers
Deepmind, G.: · 2024
Closest in time.