Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation across various domains, including medicine.
Reasoning foundations of medical diagnosis: Symbolic logic, probability, and value theory aid our understanding of how physicians reason
Robert S Ledley and Lee B Lusted · 1959
Earlier work this paper cites.
Experience with a model of sequential diagnosis
G Anthony Gorry and G Octo Barnett · 1968
Earlier work this paper cites.
Mycin: A knowledge-based computer program applied to infectious diseases
Edward H Shortliffe · 1977
Earlier work this paper cites.
Causal understanding of patient illness in medical diagnosis
Ramesh S Patil, Peter Szolovits, and William B Schwartz · 1981
Earlier work this paper cites.
Rule based expert systems: the mycin experiments of the stanford heuristic programming project (the Addison-Wesley series in artificial intelligence)
Bruce G Buchanan and Edward H Shortliffe · 1984
Earlier work this paper cites.
Toward normative expert systems: Part I the Pathfinder project
David E. Heckerman, Eric Horvitz, and Bharat N. Nathwani · 1992
Earlier work this paper cites.
Principles of mixed-initiative user interfaces
Eric Horvitz · 1999
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana · 2005
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman · 2009
Earlier work this paper cites.
Data-driven decisions for reducing readmissions for heart failure: General methodology and case study
Mohsen Bayati, Mark Braverman, Michael Gillam, Karen M Mack, George Ruiz, Mark S Smith, and Eric Horvitz · 2014
Earlier work this paper cites.
Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission
Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad · 2015
Earlier work this paper cites.
Implicit racial/ethnic bias among health care professionals and its influence on health care outcomes: a systematic review
William J Hall, Mimi V Chapman, Kent M Lee, Yesenia M Merino, Tainayah W Thomas, B Keith Payne, Eugenia Eng, Steven H Day, and Tamera Coyne-Beasley · 2015
Earlier work this paper cites.
A targeted real-time early warning score (trewscore) for septic shock
Katharine E Henry, David N Hager, Peter J Pronovost, and Suchi Saria · 2015
Earlier work this paper cites.
Patient risk stratification with time-varying parameters: a multitask learning approach
Jenna Wiens, John Guttag, and Eric Horvitz · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Dermatologist-level classification of skin cancer with deep neural networks
Andre Esteva, Brett Kuprel, Roberto A Novoa, Justin Ko, Susan M Swetter, Helen M Blau, and Sebastian Thrun · 2017
Earlier work this paper cites.
Addressing bias in machine learning algorithms: A pilot study on emotion recognition for intelligent systems
Ayanna Howard, Cha Zhang, and Eric Horvitz · 2017
Earlier work this paper cites.
Chexnet: Radiologist-level pneumonia detection on chest x-rays with deep learning
Pranav Rajpurkar, Jeremy Irvin, Kaylie Zhu, Brandon Yang, Hershel Mehta, Tony Duan, Daisy Ding, Aarti Bagul, Curtis Langlotz, Katie Shpanskaya, et al · 2017
Earlier work this paper cites.
Clinical intervention prediction and understanding with deep neural networks
Harini Suresh, Nathan Hunt, Alistair Johnson, Leo Anthony Celi, Peter Szolovits, and Marzyeh Ghassemi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A reductions approach to fair classification
Alekh Agarwal, Alina Beygelzimer, Miroslav Dudik, John Langford, and Hanna Wallach · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
Implementing machine learning in health care—addressing ethical challenges
Danton S Char, Nigam H Shah, and David Magnus · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Cited alongside, same era.
The future of the professions
Daniel Susskind and Richard Susskind · 2018
Cited alongside, same era.
Guidelines for human-AI interaction
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al · 2019
Cited alongside, same era.
Error terrain analysis for machine learning: Tool and visualizations
Rick Barraza, Russell Eames, Yan Esteve Balducci, Josh Hinds, Scott Hoogerwerf, Eric Horvitz, Ece Kamar, Jacquelyn Krones, Josh Lovejoy, Parham Mohadjer, et al · 2019
Cited alongside, same era.
Fairness and Machine Learning: Limitations and Opportunities
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2019
Cited alongside, same era.
Beyond accuracy: The role of mental models in human-AI team performance
Fairlearn: Configurable and interpretable algorithmic fairness
Ankit Kulshrestha and Ilya Safro · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Later among the works it cites.
Gender-sensitive word embeddings for healthcare
Shunit Agmon, Plia Gillis, Eric Horvitz, and Kira Radinsky · 2022
Later among the works it cites.
Prospective, multi-site study of patient outcomes after implementation of the trews machine learning-based early warning system for sepsis
Roy Adams, Katharine E Henry, Anirudh Sridharan, Hossein Soleimani, Andong Zhan, Nishi Rawat, Lauren Johnson, David N Hager, Sara E Cosgrove, Andrew Markowski, et al · 2022
Later among the works it cites.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, April 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz · 2019
Cited alongside, same era.
Updates in human-AI teams: Understanding and addressing the performance/compatibility tradeoff
Gagan Bansal, Besmira Nushi, Ece Kamar, Daniel S Weld, Walter S Lasecki, and Eric Horvitz · 2019
Cited alongside, same era.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu · 2019
Cited alongside, same era.
Interpretml: A unified framework for machine learning interpretability
Harsha Nori, Samuel Jenkins, Paul Koch, and Rich Caruana · 2019
Cited alongside, same era.
Dissecting racial bias in an algorithm used to manage the health of populations
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan · 2019
Cited alongside, same era.
Deep medicine: how artificial intelligence can make healthcare human again
Eric Topol · 2019
Cited alongside, same era.
Do no harm: a roadmap for responsible machine learning for health care
Jenna Wiens, Suchi Saria, Mark Sendak, Marzyeh Ghassemi, Vincent X Liu, Finale Doshi-Velez, Kenneth Jung, Katherine Heller, David Kale, Mohammed Saeed, et al · 2019
Cited alongside, same era.
Yuntao Bai, Andy Jones, Kamal Ndousse, et al · 2022
Later among the works it cites.
Who goes first? Influences of human-ai workflow on decision making in clinical imaging
Riccardo Fogliato, Shreya Chappidi, Matthew Lungren, Paul Fisher, Diane Wilson, Michael Fitzke, Mark Parkinson, Eric Horvitz, Kori Inkpen, and Besmira Nushi · 2022
Later among the works it cites.
Adverse events in hospitals: A quarter of medicare patients experienced harm in october 2018
Christi A Grimm · 2022
Later among the works it cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Later among the works it cites.
Holistic evaluation of language models, 2022
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, et al · 2022
Later among the works it cites.
Can large language models reason about medical questions?
Valentin Liévin, Christoffer Egeberg Hother, and Ole Winther · 2022
Later among the works it cites.
Reading between the lines: Modeling user behavior and costs in ai-assisted programming
Hussein Mozannar, Gagan Bansal, Adam Fourney, and Eric Horvitz · 2022
Later among the works it cites.
Diagnostic errors in the emergency department: A systematic review, 2022
David E Newman-Toker, Susan M Peterson, Shervin Badihian, Ahmed Hassoon, Najlla Nassery, Donna Parizadeh, Lisa M Wilson, Yuanxi Jia, Rodney Omron, Saraniya Tharmarajah, et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu · 2022
Later among the works it cites.
Impact of artificial intelligence on us medical students’ choice of radiology
Kristen Reeder and Hwan Lee · 2022
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
The safety of inpatient health care
David W Bates, David M Levine, Hojjat Salmasian, Ania Syrowatka, David M Shahian, Stuart Lipsitz, Jonathan P Zebrowski, Laura C Myers, Merranda S Logan, Christopher G Roy, et al · 2023
Closest in time.
Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models
Tiffany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al · 2023
Closest in time.
Error Analysis
Microsoft · 2023
Closest in time.
Responsible AI Mitigations and Tracker: New open-source tools for guiding mitigations in Responsible AI
Besmira Nushi and Rahee Ghosh Peshawaria · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.