Fetching the paper…
Reading the bibliography…
Recent artificial intelligence (AI) systems have reached milestones in "grand challenges" ranging from Go to protein-folding.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“Computer programs to support clinical decision making”
Edward Shortliffe · 1987
Earlier work this paper cites.
“Medicine and the computer: the promise and problems of change”
William Schwartz · 1987
Earlier work this paper cites.
“Categorical and probabilistic reasoning in medicine revisited”
Daniel Bobrow · 1994
Earlier work this paper cites.
“Free-Marginal Multirater Kappa (multirater K [free]): An Alternative to Fleiss’ Fixed-Marginal Multirater Kappa.”
Justus Randolph · 2005
Earlier work this paper cites.
“The reliability of AHRQ Common Format Harm Scales in rating patient safety events”
Tamara Williams, Marilyn Szekendi, Stephen Pavkovic, Wanda Clevenger and Julie Cerese · 2015
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Overview of the medical question answering task at TREC 2017 LiveQA.”
Asma Abacha, Eugene Agichtein, Yuval Pinter and Dina Demner-Fushman · 2017
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“PubMedQA: A dataset for biomedical research question answering”
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen and Xinghua Lu · 2019
Earlier work this paper cites.
“Bridging the Gap Between Consumers’ Medication Questions and Trusted Answers.”
Asma Abacha, Yassine Mrabet, Mark Sharp, Travis Goodwin, Sonya Shooshan and Dina Demner-Fushman · 2019
Earlier work this paper cites.
“Ethical dimensions of using artificial intelligence in health care”
Michael Rigby · 2019
Earlier work this paper cites.
“Understanding how discrimination can affect health”
David Williams, Jourdyn Lawrence, Brigette Davis and Cecilia Vu · 2019
Earlier work this paper cites.
“PubMedQA: A dataset for biomedical research question answering”
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen and Xinghua Lu · 2019
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li and Peter Liu · 2020
Earlier work this paper cites.
“Measuring massive multitask language understanding”
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song and Jacob Steinhardt · 2020
Earlier work this paper cites.
“Hidden in plain sight—reconsidering the use of race correction in clinical algorithms”
Darshali Vyas, Leo Eisenstein and David Jones · 2020
Earlier work this paper cites.
“Structural racism and health disparities: Reconfiguring the social determinants of health framework to include the root cause”
Ruqaiijah Yearby · 2020
Earlier work this paper cites.
“Measuring massive multitask language understanding”
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song and Jacob Steinhardt · 2020
Earlier work this paper cites.
“Domain-specific language model pretraining for biomedical natural language processing”
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao and Hoifung Poon · 2021
Cited alongside, same era.
“What disease does this patient have? a large-scale open domain question answering dataset from medical exams”
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang and Peter Szolovits · 2021
Cited alongside, same era.
“New creatinine-and cystatin C–based equations to estimate GFR without race”
Lesley Inker, Nwamaka Eneanya, Josef Coresh, Hocine Tighiouart, Dan Wang, Yingying Sang, Deidra Crews, Alessandro Doria, Michelle Estrella and Marc Froissart · 2021
Cited alongside, same era.
“Ethical machine learning in healthcare”
Irene Chen, Emma Pierson, Sherri Rose, Shalmali Joshi, Kadija Ferryman and Marzyeh Ghassemi · 2021
Cited alongside, same era.
“Ethical and social risks of harm from language models”
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle and Atoosa Kasirzadeh · 2021
“Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters”
boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer and Huan Sun · 2022
Later among the works it cites.
“Solving quantitative reasoning problems with language models”
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag and Theo Gutman-Solo · 2022
Later among the works it cites.
“Lamda: Language models for dialog applications”
Romal Thoppilan, Daniel De, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker and Yu Du · 2022
Later among the works it cites.
“Active Acquisition for Multimodal Temporal Data: A Challenging Decision-Making Task”
Jannik Kossen, Cătălina Cangea, Eszter Vértes, Andrew Jaegle, Viorica Patraucean, Ira Ktena, Nenad Tomasev and Danielle Belgrave · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“What disease does this patient have? a large-scale open domain question answering dataset from medical exams”
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang and Peter Szolovits · 2021
Cited alongside, same era.
“Large Language Models Encode Clinical Knowledge”
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Mahdavi, Jason Wei, Hyung Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis and Stephen Pfohl · 2022
Cited alongside, same era.
“Can large language models reason about medical questions?”
Valentin Liévin, Christoffer Hother and Ole Winther · 2022
Cited alongside, same era.
“LinkBERT: Pretraining Language Models with Document Links”
Michihiro Yasunaga, Jure Leskovec and Percy Liang · 2022
Cited alongside, same era.
“Deep bidirectional language-knowledge graph pretraining”
Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher Manning, Percy Liang and Jure Leskovec · 2022
Cited alongside, same era.
“Stanford CRFM Introduces PubMedGPT 2.7B”, https://hai.stanford.edu/news/stanford-crfm-introduces-pubmedgpt-27b , 2022
Elliot Bolton, David Hall, Michihiro Yasunaga, Tony Lee, Chris Manning and Percy Liang · 2022
Cited alongside, same era.
“BioGPT: generative pre-trained transformer for biomedical text generation and mining”
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon and Tie-Yan Liu · 2022
Cited alongside, same era.
“Holistic evaluation of language models”
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu and Ananya Kumar · 2022
Later among the works it cites.
“Red teaming language models with language models”
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese and Geoffrey Irving · 2022
Later among the works it cites.
“Chain of thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le and Denny Zhou · 2022
Later among the works it cites.
“Large Language Models Encode Clinical Knowledge”
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Mahdavi, Jason Wei, Hyung Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis and Stephen Pfohl · 2022
Later among the works it cites.
“MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering”
Ankit Pal, Logesh Umapathi and Malaikannan Sankarasubbu · 2022
Later among the works it cites.
“Capabilities of gpt-4 on medical challenge problems”
Harsha Nori, Nicholas King, Scott McKinney, Dean Carignan and Eric Horvitz · 2023
Closest in time.
“PaLM 2 Technical Report”, https://ai.google/static/documents/palm2techreport.pdf , 2023
Google · 2023
Closest in time.
“The Diagnostic and Triage Accuracy of the GPT-3 Artificial Intelligence Model”
David Levine, Rudraksh Tuwani, Benjamin Kompa, Amita Varma, Samuel Finlayson, Ateev Mehrotra and Andrew Beam · 2023
Closest in time.
“Analysis of large-language model versus human performance for genetics questions”
Dat Duong and Benjamin Solomon · 2023
Closest in time.
“ChatGPT Goes to Operating Room: Evaluating GPT-4 Performance and Its Potential in Surgical Education and Training in the Era of Large Language Models”
Namkee Oh, Gyu-Seong Choi and Woo Lee · 2023
Closest in time.
“Evaluating the performance of chatgpt in ophthalmology: An analysis of its successes and shortcomings”
Fares Antaki, Samir Touma, Daniel Milad, Jonathan El-Khoury and Renaud Duval · 2023
Closest in time.
“Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum”
John Ayers, Adam Poliak, Mark Dredze, Eric Leas, Zechariah Zhu, Jessica Kelley, Dennis Faix, Aaron Goodman, Christopher Longhurst and Michael Hogarth · 2023
Closest in time.
“Self-refine: Iterative refinement with self-feedback”
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye and Yiming Yang · 2023
Closest in time.
“DERA: Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents”
Varun Nair, Elliot Schumacher, Geoffrey Tso and Anitha Kannan · 2023
Closest in time.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Closest in time.