Fetching the paper…
Reading the bibliography…
In this work, we study the impact of QA fine-tuning data on downstream factuality.
T-rex: A large scale alignment of natural language with knowledge base triples
Elsahar, H., Vougiouklis, P., Remaci, A., Gravier, C., Hare, J., Laforest, F., and Simperl, E · 2018
Earlier work this paper cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
How can we know what language models know?, 2020
Jiang, Z., Xu, F. F., Araki, J., and Neubig, G · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Roberts, A., Raffel, C., and Shazeer, N · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision, 2022
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Impact of pretraining term frequencies on few-shot reasoning, 2022
Razeghi, Y., au2, R. L. L. I., Gardner, M., and Singh, S · 2022
Earlier work this paper cites.
Simple entity-centric questions challenge dense retrievers, 2022
Sciavolino, C., Zhong, Z., Lee, J., and Chen, D · 2022
Cited alongside, same era.
Physics of language models: Part 3.1, knowledge storage and extraction, 2023
Allen-Zhu, Z. and Li, Y · 2023
Cited alongside, same era.
Dola: Decoding by contrasting layers improves factuality in large language models, 2023
Chuang, Y.-S., Xie, Y., Luo, H., Kim, Y., Glass, J., and He, P · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models, 2023
Geva, M., Bastings, J., Filippova, K., and Globerson, A · 2023
Cited alongside, same era.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, 2023
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., and Liu, T · 2023
Cited alongside, same era.
Mistral 7b, 2023
Reinforcement learning from human feedback: Progress and challenges
Schulman, J · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation, 2023
Varshney, N., Yao, W., Zhang, H., Chen, J., and Yu, D · 2023
Later among the works it cites.
Alignment for honesty, 2023
Yang, Y., Chern, E., Qiu, X., Neubig, G., and Liu, P · 2023
Later among the works it cites.
Attention satisfies: A constraint-satisfaction lens on factual errors of language models, 2023
Yuksekgonul, M., Chandrasekaran, V., Jones, E., Gunasekar, S., Naik, R., Palangi, H., Kamar, E., and Nushi, B · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Personas as a way to model truthfulness in language models, 2023
Joshi, N., Rando, J., Saparov, A., Kim, N., and He, H · 2023
Cited alongside, same era.
Large language models struggle to learn long-tail knowledge, 2023
Kandpal, N., Deng, H., Roberts, A., Wallace, E., and Raffel, C · 2023
Cited alongside, same era.
Understanding finetuning for factual knowledge extraction from language models, 2023
Kazemi, M., Mittal, S., and Ramachandran, D · 2023
Cited alongside, same era.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories, 2023
Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., and Hajishirzi, H · 2023
Cited alongside, same era.
Inference-time intervention: Eliciting truthful answers from a language model, 2023a
Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M
Cited in the paper.
How do transformers learn topic structure: Towards a mechanistic understanding, 2023b
Li, Y., Li, Y., and Risteski, A
Cited in the paper.
Zhang, H., Diao, S., Lin, Y., Fung, Y. R., Lian, Q., Wang, X., Chen, Y., Ji, H., and Zhang, T · 2023
Later among the works it cites.
Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs
Chen, A., Shwartz-Ziv, R., Cho, K., Leavitt, M. L., and Saphra, N · 2024
Closest in time.
Does fine-tuning llms on new knowledge encourage hallucinations?, 2024
Gekhman, Z., Yona, G., Aharoni, R., Eyal, M., Feder, A., Reichart, R., and Herzig, J · 2024
Closest in time.
Unfamiliar finetuning examples control how language models hallucinate, 2024
Kang, K., Wallace, E., Tomlin, C., Kumar, A., and Levine, S · 2024
Closest in time.