Fetching the paper…
Reading the bibliography…
The widespread adoption of large language models (LLMs) across diverse AI applications is proof of the outstanding achievements obtained in several tasks, such as text mining, text generation, and question answering.
Pseudohallucinations: a conceptual history,
G. E. Berrios, T. Dening, · 1996
Earlier work this paper cites.
Mixout: Effective regularization to finetune large-scale pretrained language models,
C. Lee, K. Cho, W. Kang, · 2019
Earlier work this paper cites.
Language models are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, D. Amodei, · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale,
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, V. Stoyanov, · 2020
Earlier work this paper cites.
The curious case of hallucinations in neural machine translation,
V. Raunak, A. Menezes, M. Junczys-Dowmunt, · 2021
Earlier work this paper cites.
Evaluating attribution in dialogue systems: The begin benchmark,
N. Dziri, H. Rashkin, T. Linzen, D. Reitter, · 2021
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation,
W. Yuan, G. Neubig, P. Liu, · 2021
Earlier work this paper cites.
Emerging trends: A gentle introduction to fine-tuning,
K. W. Church, Z. Chen, Y. Ma, · 2021
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods,
S. Lin, J. Hilton, O. Evans, · 2022
Earlier work this paper cites.
On the origin of hallucinations in conversational models: Is it the datasets or the models?,
N. Dziri, S. Milton, M. Yu, O. Zaiane, S. Reddy, · 2022
Earlier work this paper cites.
Diving deep into modes of fact hallucinations in dialogue systems,
S. Das, S. Saha, R. Srihari, · 2022
Earlier work this paper cites.
FaithDial: A Faithful Benchmark for Information-Seeking Dialogue,
N. Dziri, E. Kamalloo, S. Milton, O. Zaiane, M. Yu, E. M. Ponti, S. Reddy, · 2022
Earlier work this paper cites.
Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization,
M. Cao, Y. Dong, J. Cheung, · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, R. Lowe, · 2022
Earlier work this paper cites.
Finetuned language models are zero-shot learners,
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, Q. V. Le, · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways,
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. M. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. C. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. García, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Díaz, O. Firat, M. Catasta, J. Wei, K. S. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, N. Fiedel, · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback,
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. J. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. B. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, J. Kaplan, · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models,
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. T. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, L. Zettlemoyer, · 2022
Earlier work this paper cites.
OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, H. Yang, · 2022
Earlier work this paper cites.
Let there be a clock on the beach: Reducing object hallucination in image captioning,
A. F. Biten, L. Gómez, D. Karatzas, · 2022
Earlier work this paper cites.
Language models (mostly) know what they know,
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al., · 2022
Cited alongside, same era.
Skill: structured knowledge infusion for large language models,
F. Moiseev, Z. Dong, E. Alfonseca, M. Jaggi, · 2022
Cited alongside, same era.
An efficient memory-augmented transformer for knowledge-intensive nlp tasks,
Y. Wu, Y. Zhao, B. Hu, P. Minervini, P. Stenetorp, S. Riedel, · 2022
Cited alongside, same era.
Hallucinations in large multilingual translation models,
N. M. Guerreiro, D. Alves, J. Waldendorf, B. Haddow, A. Birch, P. Colombo, A. Martins, · 2023
Cited alongside, same era.
Transformative effects of chatgpt on modern education: Emerging era of ai chatbots,
S. S. Gill, M. Xu, P. Patros, H. Wu, R. Kaur, K. Kaur, S. Fuller, M. Singh, P. Arora, A. K. Parlikad, V. Stankovski, A. Abraham, S. K. Ghosh, H. Lutfiyya, S. S. Kanhere, R. Bahsoon, O. F. Rana, S. Dustdar, R. Sakellariou, S. Uhlig, R. Buyya, · 2023
Closest in time.
Survey of hallucination in natural language generation,
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, P. Fung, · 2023
Closest in time.
Mitigating language model hallucination with interactive question-knowledge alignment,
S. Zhang, L. Pan, J. Zhao, W. Y. Wang, · 2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Dale, E. Voita, J. Lam, P. Hansanti, C. Ropers, E. Kalbassi, C. Gao, L. Barrault, M. R. Costa-jussà, · 2023
Cited alongside, same era.
mmt5: Modular multilingual pre-training solves source language hallucinations,
J. Pfeiffer, F. Piccinno, M. Nicosia, X. Wang, M. Reid, S. Ruder, · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. Gonzalez, I. C. Stoica, · 2023
Cited alongside, same era.
Evaluating correctness and faithfulness of instruction-following models for question answering,
V. Adlakha, P. BehnamGhader, X. H. Lu, N. Meade, S. Reddy, · 2023
Cited alongside, same era.
Med-halt: Medical domain hallucination test for large language models,
L. K. Umapathi, A. Pal, M. Sankarasubbu, · 2023
Cited alongside, same era.
Contrastive learning reduces hallucination in conversations,
W. Sun, Z. Shi, S. Gao, P. Ren, M. de Rijke, Z. Ren, · 2023
Cited alongside, same era.
Evaluating the factual consistency of large language models through news summarization,
D. Tam, A. Mascarenhas, S. Zhang, S. Kwan, M. Bansal, C. Raffel, · 2023
Cited alongside, same era.
“why is this misleading?”: Detecting news headline hallucinations with explanations,
J. Shen, J. Liu, D. Finnie, N. Rahmati, M. Bendersky, M. Najork, · 2023
Cited alongside, same era.
L. Pan, M. S. Saxon, W. Xu, D. Nathani, X. Wang, W. Y. Wang, · 2023
Closest in time.
Investigating the translation performance of a large multilingual language model: the case of bloom,
R. Bawden, F. Yvon, · 2023
Closest in time.
How good are gpt models at machine translation? a comprehensive evaluation,
A. Hendy, M. G. Abdelrehim, A. Sharaf, V. Raunak, M. Gabr, H. Matsushita, Y. J. Kim, M. Afify, H. H. Awadalla, · 2023
Closest in time.
Evaluating generative models for graph-to-text generation,
S. Yuan, M. Färber, · 2023
Closest in time.
Llms for knowledge graph construction and reasoning: Recent capabilities and future opportunities,
Y. Zhu, X. Wang, J. Chen, S. Qiao, Y. Ou, Y. Yao, S. Deng, H. Chen, N. Zhang, · 2023
Closest in time.
H. Liu, C. Li, Q. Wu, Y. J. Lee, · 2023
Closest in time.
Simple token-level confidence improves caption correctness,
S. Petryk, S. Whitehead, J. Gonzalez, T. Darrell, A. Rohrbach, M. Rohrbach, · 2023
Closest in time.
Album storytelling with iterative story-aware captioning and large language models,
M. Ning, Y. Xie, D. Chen, Z. Song, L. Yuan, Y. Tian, Q. Ye, L. Yuan, · 2023
Closest in time.
Fact-checking of ai-generated reports,
R. Mahmood, G. Wang, M. Kalra, P. Yan, · 2023
Closest in time.
Gptscore: Evaluate as you desire,
J. Fu, S.-K. Ng, Z. Jiang, P. Liu, · 2023
Closest in time.
The internal state of an llm knows when its lying,
A. Azaria, T. Mitchell, · 2023
Closest in time.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models,
P. Manakul, A. Liusie, M. J. Gales, · 2023
Closest in time.
N. Varshney, W. Yao, H. Zhang, J. Chen, D. Yu, · 2023
Closest in time.
Dehallucinating large language models using formal methods guided iterative prompting,
S. Jha, S. K. Jha, P. Lincoln, N. D. Bastian, A. Velasquez, S. Neema, · 2023
Closest in time.
Zero-resource hallucination prevention for large language models,
J. Luo, C. Xiao, F. Ma, · 2023
Closest in time.
Trapping llm hallucinations using tagged context prompts,
P. Feldman, J. R. Foulds, S. Pan, · 2023
Closest in time.