Fetching the paper…
Reading the bibliography…
Despite their remarkable capabilities, Large Language Models (LLMs) are prone to generate responses that contradict verifiable facts, i.e., unfaithful hallucination content.
Bertscore: Evaluating text generation with bert
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2019 · 1904
Earlier work this paper cites.
Euclidean distance mapping
Danielsson, P.-E. 1980 · 1980
Earlier work this paper cites.
Mean shift: A robust approach toward feature space analysis
Comaniciu, D.; and Meer, P. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Gaussian mixture models
Reynolds, D. A.; et al. 2009 · 2009
Earlier work this paper cites.
Bahmani, B.; Moseley, B.; Vattani, A.; Kumar, R.; and Vassilvitskii, S. 2012 · 2012
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A.; Makelov, A.; Schmidt, L.; Tsipras, D.; and Vladu, A. 2017 · 2017
Earlier work this paper cites.
Get To The Point: Summarization with Pointer-Generator Networks
See, A.; Liu, P. J.; and Manning, C. D. 2017 · 2017
Earlier work this paper cites.
Representation engineering: A top-down approach to ai transparency
Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Pan, A.; Yin, X.; Mazeika, M.; Dombrowski, A.-K.; et al. 2023 · 2017
Earlier work this paper cites.
Narayan, S.; Cohen, S. B.; and Lapata, M. 2018 · 2018
Earlier work this paper cites.
Opendialkg: Explainable conversational reasoning with attention-based walks over knowledge graphs
Moon, S.; Shah, P.; Kumar, A.; and Subba, R. 2019 · 2019
Earlier work this paper cites.
Evaluating Factuality in Generation with Dependency-level Entailment
Goyal, T.; and Durrett, G. 2020 · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
DExperts: Decoding-time controlled text generation with experts and anti-experts
Liu, A.; Sap, M.; Lu, X.; Swayamdipta, S.; Bhagavatula, C.; Smith, N. A.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Burns, C.; Ye, H.; Klein, D.; and Steinhardt, J. 2022 · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Li, X. L.; Holtzman, A.; Fried, D.; Liang, P.; Eisner, J.; Hashimoto, T.; Zettlemoyer, L.; and Lewis, M. 2022 · 2022
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2022 · 2022
Cited alongside, same era.
Adversarial Training Methods for Semi-Supervised Text Classification
Miyato, T.; Dai, A. M.; and Goodfellow, I. 2022 · 2022
Cited alongside, same era.
Introducing ChatGPT
OpenAI. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Hallucination reduction in long input text summarization
Rehman, T.; Mandal, R.; Agarwal, A.; and Sanyal, D. K. 2023 · 2023
Later among the works it cites.
Reinforcement learning from human feedback: Progress and challenges
Schulman, J. 2023 · 2023
Later among the works it cites.
Trusting your evidence: Hallucinate less with context-aware decoding
Shi, W.; Han, X.; Lewis, M.; Tsvetkov, Y.; Zettlemoyer, L.; and Yih, S. W.-t. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Later among the works it cites.
Fine-tuning language models for factuality
Tian, K.; Mitchell, E.; Yao, H.; Manning, C. D.; and Finn, C. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bang, Y.; Cahyawijaya, S.; Lee, N.; Dai, W.; Su, D.; Wilie, B.; Lovenia, H.; Ji, Z.; Yu, T.; Chung, W.; et al. 2023 · 2023
Cited alongside, same era.
Chern, I.; Chern, S.; Chen, S.; Yuan, W.; Feng, K.; Zhou, C.; He, J.; Neubig, G.; Liu, P.; et al. 2023 · 2023
Cited alongside, same era.
Dola: Decoding by contrasting layers improves factuality in large language models
Chuang, Y.-S.; Xie, Y.; Luo, H.; Kim, Y.; Glass, J.; and He, P. 2023 · 2023
Cited alongside, same era.
Diving deep into modes of fact hallucinations in dialogue systems
Das, S.; Saha, S.; and Srihari, R. K. 2023 · 2023
Cited alongside, same era.
Factkb: Generalizable factuality evaluation using language models enhanced with factual knowledge
Feng, S.; Balachandran, V.; Bai, Y.; and Tsvetkov, Y. 2023 · 2023
Cited alongside, same era.
Mixture of cluster-conditional lora experts for vision-language instruction tuning
Gou, Y.; Liu, Z.; Chen, K.; Hong, L.; Xu, H.; Li, A.; Yeung, D.-Y.; Kwok, J. T.; and Zhang, Y. 2023 · 2023
Cited alongside, same era.
Do Large Language Models Know about Facts?
Hu, X.; Chen, J.; Li, X.; Guo, Y.; Wen, L.; Yu, P. S.; and Guo, Z. 2023 · 2023
Cited alongside, same era.
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
Survey on factuality in large language models: Knowledge, retrieval and domain-specificity
Wang, C.; Liu, X.; Yue, Y.; Tang, X.; Zhang, T.; Jiayang, C.; Yao, Y.; Gao, W.; Hu, X.; Qi, Z.; et al. 2023 · 2023
Later among the works it cites.
Baichuan 2: Open large-scale language models
Yang, A.; Xiao, B.; Wang, B.; Zhang, B.; Bian, C.; Yin, C.; Lv, C.; Pan, D.; Wang, D.; Yan, D.; et al. 2023 · 2023
Later among the works it cites.
Why Does ChatGPT Fall Short in Providing Truthful Answers?
Zheng, S.; Huang, J.; and Chang, K. C.-C. 2023 · 2023
Later among the works it cites.
Truth forest: Toward multi-scale truthfulness in large language models through intervention without tuning
Chen, Z.; Sun, X.; Jiao, X.; Lian, F.; Kang, Z.; Wang, D.; and Xu, C. 2024 · 2024
Closest in time.
SH2: Self-Highlighted Hesitation Helps You Decode More Truthfully
Kai, J.; Zhang, T.; Hu, H.; and Lin, Z. 2024 · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Li, K.; Patel, O.; Viégas, F.; Pfister, H.; and Wattenberg, M. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2024 · 2024
Closest in time.
Chain-of-thought reasoning without prompting
Wang, X.; and Zhou, D. 2024 · 2024
Closest in time.
PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications
Yang, D.; Wei, J.; Xiao, D.; Wang, S.; Wu, T.; Li, G.; Li, M.; Wang, S.; Chen, J.; Jiang, Y.; et al. 2024 · 2024
Closest in time.
Truthx: Alleviating hallucinations by editing large language models in truthful space
Zhang, S.; Yu, T.; and Feng, Y. 2024 · 2024
Closest in time.
Lima: Less is more for alignment
Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; et al. 2024 · 2024
Closest in time.