Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are trained on web-scale corpora that inevitably include contradictory factual information from sources of varying reliability.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Multifc: A real-world multi-domain dataset for evidence-based fact checking of claims
Augenstein, I., Lioma, C., Wang, D., Lima, L. C., Hansen, C., Hansen, C., and Simonsen, J. G. (2019) · 1909
Earlier work this paper cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S. (2019) · 1909
Earlier work this paper cites.
Defeasible reasoning
Pollock, J. L. (1987) · 1987
Earlier work this paper cites.
Sic transit gloria telae: Towards an understanding of the web’s decay
Bar-Yossef, Z., Broder, A. Z., Kumar, R., and Tomkins, A. (2004) · 2004
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
Wadden, D., Lin, S., Lo, K., Wang, L. L., van Zuylen, M., Cohan, A., and Hajishirzi, H. (2020) · 2004
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Maynez, J., Narayan, S., Bohnet, B., and McDonald, R. (2020) · 2005
Earlier work this paper cites.
Facts as experts: Adaptable and interpretable neural memory over symbolic knowledge
Verga, P., Sun, H., Soares, L. B., and Cohen, W. W. (2020) · 2007
Earlier work this paper cites.
Multi-fact correction in abstractive text summarization
Dong, Y., Wang, S., Gan, Z., Cheng, Y., Cheung, J. C. K., and Liu, J. (2020) · 2010
Earlier work this paper cites.
" liar, liar pants on fire": A new benchmark dataset for fake news detection
Wang, W. Y. (2017) · 2017
Earlier work this paper cites.
T-rex: A large scale alignment of natural language with knowledge base triples
Elsahar, H., Vougiouklis, P., Remaci, A., Gravier, C., Hare, J., Laforest, F., and Simperl, E. (2018) · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A. (2018) · 2018
Earlier work this paper cites.
Deepcopy: Grounded response generation with hierarchical pointer networks
Yavuz, S., Rastogi, A., Chao, G.-L., and Hakkani-Tur, D. (2019) · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. (2020) · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Cited alongside, same era.
Towards faithful neural table-to-text generation with content-matching constraints
Wang, Z., Wang, X., An, B., Yu, D., and Chen, C. (2020) · 2020
Cited alongside, same era.
Truthful ai: Developing and governing ai that does not lie
Evans, O., Cotton-Barratt, O., Finnveden, L., Bales, A., Balwit, A., Wills, P., Righetti, L., and Saunders, W. (2021) · 2021
Cited alongside, same era.
X-fact: A new benchmark dataset for multilingual fact checking
Gupta, A. and Srikumar, V. (2021) · 2021
Cited alongside, same era.
Evaluating the factual consistency of large language models through summarization
Tam, D., Mascarenhas, A., Zhang, S., Kwan, S., Bansal, M., and Raffel, C. (2022) · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. (2022) · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. (2022) · 2022
Later among the works it cites.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, S., Hilton, J., and Evans, O. (2021) · 2021
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G. (2022) · 2022
Cited alongside, same era.
A survey on automated fact-checking
Guo, Z., Schlichtkrull, M., and Vlachos, A. (2022) · 2022
Cited alongside, same era.
Kamel: Knowledge analysis with multitoken entities in language models
Kalo, J.-C. and Fichtel, L. (2022) · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. (2022) · 2022
Cited alongside, same era.
Firebolt: Weak supervision under weaker assumptions
Kuang, Z., Arachie, C. G., Liang, B., Narayana, P., DeSalvo, G., Quinn, M. S., Huang, B., Downs, G., and Yang, Y. (2022) · 2022
Cited alongside, same era.
Internet-augmented language models through few-shot prompting for open-domain question answering
Lazaridou, A., Gribovskaya, E., Stokowiec, W., and Grigorev, N. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. (2023) · 2023
Closest in time.
Artificial intelligence is ineffective and potentially harmful for fact checking
DeVerna, M. R., Yan, H. Y., Yang, K.-C., and Menczer, F. (2023) · 2023
Closest in time.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. (2023) · 2023
Closest in time.
Boardgameqa: A dataset for natural language reasoning with contradictory information
Kazemi, M., Yuan, Q., Bhatia, D., Kim, N., Xu, X., Imbrasaite, V., and Ramachandran, D. (2023) · 2023
Closest in time.
Whose opinions do language models reflect?
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T. (2023) · 2023
Closest in time.
Sun, K., Xu, Y. E., Zha, H., Liu, Y., and Dong, X. L. (2023) · 2023
Closest in time.
Do llms exhibit human-like response biases? a case study in survey design
Tjuatja, L., Chen, V., Wu, S. T., Talwalkar, A., and Neubig, G. (2023) · 2023
Closest in time.
Simple synthetic data reduces sycophancy in large language models
Wei, J., Huang, D., Lu, Y., Zhou, D., and Le, Q. V. (2023) · 2023
Closest in time.
Large language models as optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X. (2023) · 2023
Closest in time.