Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, yet they often struggle with maintaining factual accuracy, particularly in knowledge-intensive domains like healthcare.
An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition
Tsatsaronis, G.; Balikas, G.; Malakasiotis, P.; Partalas, I.; Zschunke, M.; Alvers, M. R.; Weissenborn, D.; Krithara, A.; Petridis, S.; Polychronopoulos, D.; et al. 2015 · 2015
Earlier work this paper cites.
PubMedQA: A Dataset for Biomedical Research Question Answering
Jin, Q.; Dhingra, B.; Liu, Z.; Cohen, W.; and Lu, X. 2019 · 2019
Earlier work this paper cites.
Defending against neural fake news
Zellers, R.; Holtzman, A.; Rashkin, H.; Bisk, Y.; Farhadi, A.; Roesner, F.; and Choi, Y. 2019 · 2019
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2020
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert
Khattab, O.; and Zaharia, M. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; et al. 2020 · 2020
Earlier work this paper cites.
GPT-3, Bloviator: OpenAI’s language generator has no idea what it’s talking about
Marcus, G.; and Davis, E. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Yih, S. 2020 · 2020
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; and Szolovits, P. 2021 · 2021
Earlier work this paper cites.
KILT: a Benchmark for Knowledge Intensive Language Tasks
Petroni, F.; Piktus, A.; Fan, A.; Lewis, P.; Yazdani, M.; Cao, N.; Thorne, J.; Jernite, Y.; Karpukhin, V.; Maillard, J.; et al. 2021 · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
Borgeaud, S.; Mensch, A.; Hoffmann, J.; Cai, T.; Rutherford, E.; Millican, K.; Van Den Driessche, G. B.; Lespiau, J.-B.; Damoc, B.; Clark, A.; et al. 2022 · 2022
Earlier work this paper cites.
Bioreader: a retrieval-enhanced text-to-text transformer for biomedical literature
Frisoni, G.; Mizutani, M.; Moro, G.; and Valgimigli, L. 2022 · 2022
Cited alongside, same era.
Literature-Augmented Clinical Outcome Prediction
Naik, A.; Parasa, S.; Feldman, S.; Wang, L. L.; and Hope, T. 2022 · 2022
Cited alongside, same era.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A.; Umapathi, L. K.; and Sankarasubbu, M. 2022 · 2022
Cited alongside, same era.
Bang, Y.; Cahyawijaya, S.; Lee, N.; Dai, W.; Su, D.; Wilie, B.; Lovenia, H.; Ji, Z.; Yu, T.; Chung, W.; et al. 2023 · 2023
Cited alongside, same era.
Chern, I.-C.; Chern, S.; Chen, S.; Yuan, W.; Feng, K.; Zhou, C.; He, J.; Neubig, G.; and Liu, P. 2023 · 2023
Fine-tuning Language Models for Factuality
Tian, K.; Mitchell, E.; Yao, H.; Manning, C.; and Finn, C. 2023 · 2023
Later among the works it cites.
Factcheck-GPT: End-to-End Fine-Grained Document-Level Fact-Checking and Correction of LLM Output
Wang, Y.; Gangi Reddy, R.; Mujahid, Z. M.; Arora, A.; Rubashevskii, A.; Geng, J.; Afzal, O. M.; Pan, L.; Borenstein, N.; Pillai, A.; et al. 2023 · 2023
Later among the works it cites.
Relic: Investigating large language model responses using self-consistency
Cheng, F.; Zouhar, V.; Arora, S.; Sachan, M.; Strobelt, H.; and El-Assady, M. 2024 · 2024
Closest in time.
Language Models Hallucinate, but May Excel at Fact Verification
Guan, J.; Dodge, J.; Wadden, D.; Huang, M.; and Peng, H. 2024 · 2024
Closest in time.
SimPO: Simple Preference Optimization with a Reference-Free Reward
Meng, Y.; Xia, M.; and Chen, D. 2024 · 2024
Closest in time.
Capabilities of Gemini Models in Medicine
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Retrieval-Augmented Generation for Large Language Models: A Survey
Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; and Wang, H. 2023 · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y. J.; Madotto, A.; and Fung, P. 2023 · 2023
Cited alongside, same era.
Active Retrieval Augmented Generation
Jiang, Z.; Xu, F. F.; Gao, L.; Sun, Z.; Liu, Q.; Dwivedi-Yu, J.; Yang, Y.; Callan, J.; and Neubig, G. 2023 · 2023
Cited alongside, same era.
Retrieve, summarize, and verify: how will ChatGPT affect information seeking from the medical literature?
Jin, Q.; Leaman, R.; and Lu, Z. 2023 · 2023
Cited alongside, same era.
In-context retrieval-augmented language models
Ram, O.; Levine, Y.; Dalmedigos, I.; Muhlgay, D.; Shashua, A.; Leyton-Brown, K.; and Shoham, Y. 2023 · 2023
Cited alongside, same era.
Saab, K.; Tu, T.; Weng, W.-H.; Tanno, R.; Stutz, D.; Wulczyn, E.; Zhang, F.; Strother, T.; Park, C.; Vedadi, E.; et al. 2024 · 2024
Closest in time.
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Wang, H.; Xiong, W.; Xie, T.; Zhao, H.; and Zhang, T. 2024 · 2024
Closest in time.
Long-form factuality in large language models
Wei, J.; Yang, C.; Song, X.; Lu, Y.; Hu, N.; Tran, D.; Peng, D.; Liu, R.; Huang, D.; Du, C.; et al. 2024 · 2024
Closest in time.
Benchmarking Retrieval-Augmented Generation for Medicine
Xiong, G.; Jin, Q.; Lu, Z.; and Zhang, A. 2024 · 2024
Closest in time.
Qwen2 Technical Report
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; et al. 2024 · 2024
Closest in time.