Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but the quality bar for medical and clinical applications is high.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation”
Kishore Papineni, Salim Roukos, Todd Ward and Wei-Jing Zhu · 2002
Earlier work this paper cites.
“Health literacy interventions and outcomes: an updated systematic review.”
Nancy Berkman, Stacey Sheridan, Katrina Donahue, David Halpern, Anthony Viera, Karen Crotty, Audrey Holland, Michelle Brasure, Kathleen Lohr and Elizabeth Harden · 2011
Earlier work this paper cites.
“Development of the Patient Education Materials Assessment Tool (PEMAT): a new measure of understandability and actionability for print and audiovisual patient information”
Sarah Shoemaker, Michael Wolf and Cindy Brach · 2014
Earlier work this paper cites.
“An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition”
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis and Dimitris Polychronopoulos · 2015
Earlier work this paper cites.
“The reliability of AHRQ Common Format Harm Scales in rating patient safety events”
Tamara Williams, Marilyn Szekendi, Stephen Pavkovic, Wanda Clevenger and Julie Cerese · 2015
Earlier work this paper cites.
“Overview of the medical question answering task at TREC 2017 LiveQA.”
Asma Abacha, Eugene Agichtein, Yuval Pinter and Dina Demner-Fushman · 2017
Earlier work this paper cites.
“TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension”
Mandar Joshi, Eunsol Choi, Daniel Weld and Luke Zettlemoyer · 2017
Earlier work this paper cites.
“Decoupled weight decay regularization”
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
“Scale development: ten main limitations and recommendations to improve future research practices”
Fabiane Morgado, Juliana Meireles, Clara Neves, Ana Amaral and Maria Ferreira · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Measuring harm in healthcare: optimizing adverse event review”
Kathleen Walsh, Polina Harik, Kathleen Mazor, Deborah Perfetto, Milena Anatchkova, Colleen Biggins, Joann Wagner, Pamela Schoettker, Cassandra Firneno and Robert Klugman · 2017
Earlier work this paper cites.
“Best practices for developing and validating scales for health, social, and behavioral research: a primer”
Godfred Boateng, Torsten Neilands, Edward Frongillo, Hugo Melgar-Quiñonez and Sera Young · 2018
Earlier work this paper cites.
“emrqa: A large corpus for question answering on electronic medical records”
Anusri Pampari, Preethi Raghavan, Jennifer Liang and Jian Peng · 2018
Earlier work this paper cites.
“Bridging the Gap Between Consumers’ Medication Questions and Trusted Answers.”
Asma Abacha, Yassine Mrabet, Mark Sharp, Travis Goodwin, Sonya Shooshan and Dina Demner-Fushman · 2019
Earlier work this paper cites.
“SciBERT: A pretrained language model for scientific text”
Iz Beltagy, Kyle Lo and Arman Cohan · 2019
Earlier work this paper cites.
“Counterfactual fairness in text classification through robustness”
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed Chi and Alex Beutel · 2019
Earlier work this paper cites.
“PubMedQA: A dataset for biomedical research question answering”
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen and Xinghua Lu · 2019
Earlier work this paper cites.
“Model cards for model reporting”
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Raji and Timnit Gebru · 2019
Earlier work this paper cites.
“Perturbation sensitivity analysis to detect unintended model biases”
Vinodkumar Prabhakaran, Ben Hutchinson and Margaret Mitchell · 2019
Earlier work this paper cites.
“TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages”
Jonathan Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev and Jennimaria Palomaki · 2020
Earlier work this paper cites.
“Measuring massive multitask language understanding”
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song and Jacob Steinhardt · 2020
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei · 2020
Earlier work this paper cites.
“BioBERT: a pre-trained biomedical language representation model for biomedical text mining”
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan So and Jaewoo Kang · 2020
Earlier work this paper cites.
“Pretrained language models for biomedical and clinical tasks: Understanding and extending the state-of-the-art”
Patrick Lewis, Myle Ott, Jingfei Du and Veselin Stoyanov · 2020
Earlier work this paper cites.
“DARE: Data augmented relation extraction with gpt-2”
Yannis Papanikolaou and Andrea Pierleoni · 2020
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer.”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li and Peter Liu · 2020
Earlier work this paper cites.
“Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing”
Inioluwa Raji, Andrew Smart, Rebecca White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron and Parker Barnes · 2020
Earlier work this paper cites.
“Expert discussions improve comprehension of difficult cases in medical image assessment”
Mike Schaekermann, Carrie Cai, Abigail Huang and Rory Sayres · 2020
Earlier work this paper cites.
“BioMegatron: Larger biomedical domain language model”
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi and Raghav Mani · 2020
Earlier work this paper cites.
“Hidden in plain sight—reconsidering the use of race correction in clinical algorithms”
Darshali Vyas, Leo Eisenstein and David Jones · 2020
Earlier work this paper cites.
“Predicting conversion to wet age-related macular degeneration using deep learning”
Jason Yim, Reena Chopra, Terry Spitz, Jim Winkens, Annette Obika, Christopher Kelly, Harry Askham, Marko Lukic, Josef Huemer and Katrin Fasler · 2020
Earlier work this paper cites.
“Hurtful words: quantifying biases in clinical contextual word embeddings”
Haoran Zhang, Amy Lu, Mohamed Abdalla, Matthew McDermott and Marzyeh Ghassemi · 2020
Earlier work this paper cites.
“GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow” If you use this software, please cite it using these metadata
Sid Black, Leo Gao, Phil Wang, Connor Leahy and Stella Biderman · 2021
Earlier work this paper cites.
“On the opportunities and risks of foundation models”
Rishi Bommasani, Drew Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael Bernstein, Jeannette Bohg, Antoine Bosselut and Emma Brunskill · 2021
Cited alongside, same era.
“Ethical machine learning in healthcare”
Irene Chen, Emma Pierson, Sherri Rose, Shalmali Joshi, Kadija Ferryman and Marzyeh Ghassemi · 2021
Cited alongside, same era.
“Training verifiers to solve math word problems”
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse and John Schulman · 2021
Cited alongside, same era.
“Deep learning-enabled medical computer vision”
Andre Esteva, Katherine Chou, Serena Yeung, Nikhil Naik, Ali Madani, Ali Mottaghi, Yun Liu, Eric Topol, Jeff Dean and Richard Socher · 2021
Cited alongside, same era.
“Datasheets for datasets”
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Vaughan, Hanna Wallach, Halé Iii and Kate Crawford · 2021
Cited alongside, same era.
“Ptr: Prompt tuning with rules for text classification”
Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu and Maosong Sun · 2022
Closest in time.
“Training Compute-Optimal Large Language Models”
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego Casas, Lisa Hendricks, Johannes Welbl and Aidan Clark · 2022
Closest in time.
“ScholarBERT: Bigger is Not Always Better”
Zhi Hong, Aswathy Ajith, Gregory Pauloski, Eamon Duede, Carl Malamud, Roger Magoulas, Kyle Chard and Ian Foster · 2022
Closest in time.
“Language models (mostly) know what they know”
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Dodds, Nova DasSarma and Eli Tran-Johnson · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Domain-specific language model pretraining for biomedical natural language processing”
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao and Hoifung Poon · 2021
Cited alongside, same era.
“Ppt: Pre-trained prompt tuning for few-shot learning”
Yuxian Gu, Xu Han, Zhiyuan Liu and Minlie Huang · 2021
Cited alongside, same era.
“Ethics and governance of artificial intelligence for health”
WHO Guidance · 2021
Cited alongside, same era.
“Moving beyond “algorithmic bias is a data problem””
Sara Hooker · 2021
Cited alongside, same era.
“What disease does this patient have? a large-scale open domain question answering dataset from medical exams”
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang and Peter Szolovits · 2021
Cited alongside, same era.
“Identifying credible sources of health information in social media: Principles and attributes”
Raynard Kington, Stacey Arnesen, Wen-Ying Chou, Susan Curry, David Lazer and Antonia Villarruel · 2021
Cited alongside, same era.
“Algorithmic monoculture and social welfare”
Jon Kleinberg and Manish Raghavan · 2021
Cited alongside, same era.
Takeshi Kojima, Shixiang Gu, Machel Reid, Yutaka Matsuo and Yusuke Iwasawa · 2022
Closest in time.
“Rethinking Explainability as a Dialogue: A Practitioner’s Perspective”
Himabindu Lakkaraju, Dylan Slack, Yuxin Chen, Chenhao Tan and Sameer Singh · 2022
Closest in time.
“Can language models learn from explanations in context?”
Andrew Lampinen, Ishita Dasgupta, Stephanie Chan, Kory Matthewson, Michael Tessler, Antonia Creswell, James McClelland, Jane Wang and Felix Hill · 2022
Closest in time.
“Solving quantitative reasoning problems with language models”
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag and Theo Gutman-Solo · 2022
Closest in time.
“Holistic evaluation of language models”
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu and Ananya Kumar · 2022
Closest in time.
“Can large language models reason about medical questions?”
Valentin Liévin, Christoffer Hother and Ole Winther · 2022
Closest in time.
“Teaching Models to Express Their Uncertainty in Words”
Stephanie Lin, Jacob Hilton and Owain Evans · 2022
Closest in time.
“The medical algorithmic audit”
Xiaoxuan Liu, Ben Glocker, Melissa McCradden, Marzyeh Ghassemi, Alastair Denniston and Lauren Oakden-Rayner · 2022
Closest in time.
“BioGPT: generative pre-trained transformer for biomedical text generation and mining”
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon and Tie-Yan Liu · 2022
Closest in time.
“Artificial Intelligence in Health Care: The Hope, the Hype, the Promise, the Peril”
Michael Matheny, Sonoo Israni, Mahnoor Ahmed and Danielle Whicher · 2022
Closest in time.
“The Blueprint for an AI Bill of Rights: Making Automated Systems Work for the American People”, https://www.whitehouse.gov/wp-content/uploads/2022/10/Blueprint-for-an-AI-Bill-of-Rights.pdf , 2022
White of Science and Technology Policy · 2022
Closest in time.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama and Alex Ray · 2022
Closest in time.
“MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering”
Ankit Pal, Logesh Umapathi and Malaikannan Sankarasubbu · 2022
Closest in time.
“Healthsheet: Development of a Transparency Artifact for Health Datasets”
Negar Rostamzadeh, Diana Mincu, Subhrajit Roy, Andrew Smart, Lauren Wilcox, Mahima Pushkarna, Jessica Schrouff, Razvan Amironesei, Nyalleng Moorosi and Katherine Heller · 2022
Closest in time.
“BLOOM: A 176B-Parameter Open-Access Multilingual Language Model”
Teven Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Luccioni, François Yvon and Matthias Gallé · 2022
Closest in time.
“Operationalizing and Implementing Pretrained, Large Artificial Intelligence Linguistic Models in the US Health Care System: Outlook of Generative Pretrained Transformer 3 (GPT-3) as a Service Model”
Emre Sezgin, Joseph Sirrianni and Simon Linwood · 2022
Closest in time.
“Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models”
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Shoeb, Abubakar Abid, Adam Fisch, Adam Brown, Adam Santoro, Aditya Gupta and Adrià Garriga-Alonso · 2022
Closest in time.
“Galactica: A Large Language Model for Science”
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez and Robert Stojnic · 2022
Closest in time.
“Lamda: Language models for dialog applications”
Romal Thoppilan, Daniel De, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker and Yu Du · 2022
Closest in time.
“Plex: Towards reliability using pretrained large model extensions”
Dustin Tran, Jeremiah Liu, Michael Dusenberry, Du Phan, Mark Collier, Jie Ren, Kehang Han, Zi Wang, Zelda Mariet and Huiyi Hu · 2022
Closest in time.
“Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters”
boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer and Huan Sun · 2022
Closest in time.
“Self-consistency improves chain of thought reasoning in language models”
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi and Denny Zhou · 2022
Closest in time.
“Emergent abilities of large language models”
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou and Donald Metzler · 2022
Closest in time.
“Chain of thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le and Denny Zhou · 2022
Closest in time.
“Deep bidirectional language-knowledge graph pretraining”
Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher Manning, Percy Liang and Jure Leskovec · 2022
Closest in time.
“LinkBERT: Pretraining Language Models with Document Links”
Michihiro Yasunaga, Jure Leskovec and Percy Liang · 2022
Closest in time.
“Retrieval of Soft Prompt Enhances Zero-Shot Task Generalization”
Seonghyeon Ye, Joel Jang, Doyoung Kim, Yongrae Jo and Minjoon Seo · 2022
Closest in time.
“OPT: Open pre-trained transformer language models”
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li and Xi Lin · 2022
Closest in time.
“Least-to-Most Prompting Enables Complex Reasoning in Large Language Models”
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le and Ed Chi · 2022
Closest in time.