Fetching the paper…
Reading the bibliography…
There is vivid research on adapting Large Language Models (LLMs) to perform a variety of tasks in high-stakes domains such as healthcare.
Taxonomy of education objectives Book 1-Cognitive domain
Benjamin S Bloom. 1956 · 1956
Earlier work this paper cites.
The role of knowledge in discourse comprehension: A construction-integration model
Walter Kintsch. 1988 · 1988
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Ye Liu, Shaika Chowdhury, Chenwei Zhang, Cornelia Caragea, and Philip S Yu. 2020 · 2008
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2009
Earlier work this paper cites.
Index expansion for machine reading and question answering
Giuseppe Attardi, Luca Atzori, Maria Simi, et al. 2012 · 2012
Earlier work this paper cites.
Machine reading of biomedical texts about alzheimers disease
Roser Morante, Martin Krallinger, Alfonso Valencia, and Walter Daelemans. 2012 · 2012
Earlier work this paper cites.
Modeling biological processes for reading comprehension
Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Bloom’s taxonomy of cognitive learning objectives
Nancy E Adams. 2015 · 2015
Earlier work this paper cites.
An overview of the bioasq large-scale biomedical semantic indexing and question answering competition
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, Yannis Almirantis, John Pavlopoulos, Nicolas Baskiotis, Patrick Gallinari, Thierry Artieres, Axel Ngonga, Norman Heino, Eric Gaussier, Liliana Barrio-Alvers, Michael Schroeder, Ion Androutsopoulos, and Georgios Paliouras. 2015 · 2015
Earlier work this paper cites.
Overview of the medical question answering task at TREC 2017 LiveQA
Asma Ben Abacha, Eugene Agichtein, Yuval Pinter, and Dina Demner-Fushman. 2017 · 2017
Earlier work this paper cites.
Towards A Rigorous Science of Interpretable Machine Learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer. 2018 · 2018
Earlier work this paper cites.
Question answering as global reasoning over semantic abstractions
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Dan Roth. 2018 · 2018
Earlier work this paper cites.
A question-entailment approach to question answering
Asma Ben Abacha and Dina Demner-Fushman. 2019 · 2019
Earlier work this paper cites.
Bridging the gap between consumers’ medication questions and trusted answers
Asma Ben Abacha, Yassine Mrabet, Mark Sharp, Travis Goodwin, Sonya E. Shooshan, and Dina Demner-Fushman. 2019 · 2019
Earlier work this paper cites.
Evaluating question answering evaluation
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Comprehensive Multi-Dataset Evaluation of Reading Comprehension
Dheeru Dua, Ananth Gottumukkala, Alon Talmor, Matt Gardner, and Sameer Singh. 2019 · 2019
Earlier work this paper cites.
MRQA 2019 Shared Task: Evaluating Generalization in Reading Comprehension
Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019 · 2019
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019 · 2019
Earlier work this paper cites.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019 · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Cited alongside, same era.
MultiQA: An Empirical Investigation of Generalization and Transfer in Reading Comprehension
Alon Talmor and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
HEAD-QA: A healthcare dataset for complex reasoning
David Vilares and Carlos Gómez-Rodríguez. 2019 · 2019
Cited alongside, same era.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020 · 2020
Cited alongside, same era.
BioMRC: A dataset for biomedical machine reading comprehension
Dimitris Pappas, Petros Stavropoulos, Ion Androutsopoulos, and Ryan McDonald. 2020 · 2020
Cited alongside, same era.
Emergent Abilities of Large Language Models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022 · 2022
Later among the works it cites.
Falcon-40B: an open large language model with state-of-the-art performance
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023 · 2023
Later among the works it cites.
Overview of the MEDIQA-chat 2023 shared tasks on the summarization & generation of doctor-patient conversations
Asma Ben Abacha, Wen-wai Yim, Griffin Adams, Neal Snider, and Meliha Yetisgen. 2023a · 2023
Later among the works it cites.
Redpajama-data: An open source recipe to reproduce llama training dataset
Together Computer. 2023 · 2023
Later among the works it cites.
Free dolly: Introducing the world’s first truly open instruction-tuned llm
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Question-driven summarization of answers to consumer health questions
Max Savery, Asma Ben Abacha, Soumya Gayen, and Dina Demner-Fushman. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
Question answering with long multiple-span answers
Ming Zhu, Aman Ahuja, Da-Cheng Juan, Wei Wei, and Chandan K. Reddy. 2020 · 2020
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021 · 2021
Cited alongside, same era.
The factual inconsistency problem in abstractive text summarization: A survey
Yichong Huang, Xiachong Feng, Xiaocheng Feng, and Bing Qin. 2021 · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021 · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021 · 2021
Cited alongside, same era.
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023 · 2023
Later among the works it cites.
QLoRA: Efficient Finetuning of Quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Later among the works it cites.
Medalign: A clinician-generated dataset for instruction following with electronic medical records
Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo P. Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak, Birju S. Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay S. Chaudhari, Nima Aghaeepour, Christopher Sharp, Michael A. Pfeffer, Percy Liang, Jonathan H. Chen, Keith E. Morse, Emma P. Brunskill, Jason A. Fries, and Nigam H. Shah. 2023 · 2023
Later among the works it cites.
Medalpaca–an open-source collection of medical conversational ai models and training data
Tianyu Han, Lisa C Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K Bressem. 2023 · 2023
Later among the works it cites.
Medeval: A multi-level, multi-task, and multi-domain medical benchmark for language model evaluation
Zexue He, Yu Wang, An Yan, Yao Liu, Eric Y Chang, Amilcare Gentili, Julian McAuley, and Chun-Nan Hsu. 2023 · 2023
Later among the works it cites.
Bioasq-qa: A manually curated corpus for biomedical question answering
Anastasia Krithara, Anastasios Nentidis, Konstantinos Bougiatiotis, and Georgios Paliouras. 2023 · 2023
Later among the works it cites.
Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge
Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. 2023 · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms
NLP Team MosaicML. 2023 · 2023
Later among the works it cites.
Can generalist foundation models outcompete special-purpose tuning? case study in medicine
Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, Renqian Luo, Scott Mayer McKinney, Robert Osazuwa Ness, Hoifung Poon, Tao Qin, Naoto Usuyama, Chris White, and Eric Horvitz. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Later among the works it cites.
Comparative analysis of large language models in the royal college of ophthalmologists fellowship exams
Raffaele Raimondi, Nikolaos Tzoumas, Thomas Salisbury, Sandro Di Simplicio, and Mario R Romano. 2023 · 2023
Later among the works it cites.
Clinical camel: An open expert-level medical language model with dialogue-based knowledge encoding
Augustin Toma, Patrick R. Lawler, Jimmy Ba, Rahul G. Krishnan, Barry B. Rubin, and Bo Wang. 2023 · 2023
Later among the works it cites.
Med-halt: Medical domain hallucination test for large language models
Logesh Kumar Umapathi, Ankit Pal, and Malaikannan Sankarasubbu. 2023 · 2023
Later among the works it cites.
Clinical text summarization: Adapting large language models can outperform human experts
Dave Van Veen, Cara Van Uden, Louis Blankemeier, Jean-Benoit Delbrouck, Asad Aali, Christian Bluethgen, Anuj Pareek, Malgorzata Polacin, William Collins, Neera Ahuja, Curtis P. Langlotz, Jason Hom, Sergios Gatidis, John Pauly, and Akshay S. Chaudhari. 2023 · 2023
Later among the works it cites.
Pmc-llama: Towards building open-source language models for medicine
Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023 · 2023
Later among the works it cites.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023 · 2023
Later among the works it cites.