Fetching the paper…
Reading the bibliography…
Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reasoning with factual recall.
Medical problem solving: An analysis of clinical reasoning
Elstein, A. S., Shulman, L. S., and Sprafka, S. A · 1978
Earlier work this paper cites.
Knowledge based solution strategies in medical reasoning
Patel, V. L. and Groen, G. J · 1986
Earlier work this paper cites.
The clinical reasoning process
Barrows, H. S. and Feltovich, P. J · 1987
Earlier work this paper cites.
Clinical reasoning, decisionmaking, and action: Thinking critically and clinically
Benner, P., Hughes, R. G., and Sutphen, M · 2008
Earlier work this paper cites.
A universal model of diagnostic reasoning
Croskerry, P · 2009
Earlier work this paper cites.
Errors in clinical reasoning: causes and remedial strategies
Scott, I. A · 2009
Earlier work this paper cites.
An integrated model of clinical reasoning: dual-process theory of cognition and metacognition
Marcum, J. A · 2012
Earlier work this paper cites.
Scripts, plans, goals, and understanding: An inquiry into human knowledge structures
Schank, R. C. and Abelson, R. P · 2013
Earlier work this paper cites.
Using script theory to cultivate illness script formation and clinical reasoning in health professions education
Lubarsky, S., Dory, V., Audétat, M.-C., Custers, E., and Charlin, B · 2015
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X · 2019
Earlier work this paper cites.
Reasoning processes in clinical reasoning: from the perspective of cognitive psychology
Shin, H. S · 2019
Earlier work this paper cites.
Head-qa: A healthcare dataset for complex reasoning
Vilares, D. and Gómez-Rodríguez, C · 2019
Earlier work this paper cites.
Five decades of research and theorization on clinical reasoning: a critical review
Yazdani, S. and Hoseini Abardeh, M · 2019
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., and Poon, H · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P · 2021
Earlier work this paper cites.
Embracing complexity with systems thinking in general practitioners’ clinical reasoning helps handling uncertainty
Stolper, E., Van Royen, P., Jack, E., Uleman, J., and Olde Rikkert, M · 2021
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., Umapathi, L. K., and Sankarasubbu, M · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Generating sequences by learning to self-correct
Welleck, S., Lu, X., West, P., Brahman, F., Shen, T., Khashabi, D., and Choi, Y · 2022
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Cited alongside, same era.
Backtracking improves generation safety
Zhang, Y., Chi, J., Nguyen, H., Upasani, K., Bikel, D. M., Weston, J., and Smith, E. M · 2024
Later among the works it cites.
Evaluation of openai o1: Opportunities and challenges of agi
Zhong, T., Liu, Z., Pan, Y., Zhang, Y., Zhou, Y., Liang, S., Wu, Z., Lyu, Y., Shu, P., Yu, X., et al · 2024
Later among the works it cites.
Lessons from red teaming 100 generative ai products
Bullwinkel, B., Minnich, A., Chawla, S., Lopez, G., Pouliot, M., Maxwell, W., de Gruyter, J., Pratt, K., Qi, S., Chikanov, N., et al · 2025
Closest in time.
Red teaming chatgpt in medicine to yield real-world insights on model behavior
Chang, C. T., Farah, H., Gui, H., Rezaei, S. J., Bou-Khalil, C., Park, Y.-J., Swaminathan, A., Omiye, J. A., Kolluri, A., Chaurasia, A., et al · 2025
Closest in time.
Sft or rl? an early investigation into training r1-like reasoning large vision-language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Navigating the grey area: How expressions of uncertainty and overconfidence affect language models
Zhou, K., Jurafsky, D., and Hashimoto, T · 2023
Cited alongside, same era.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al · 2024
Cited alongside, same era.
Medical large language models are susceptible to targeted misinformation attacks
Han, T., Nebelung, S., Khader, F., Wang, T., Müller-Franzes, G., Kuhl, C., Försch, S., Kleesiek, J., Haarburger, C., Bressem, K. K., et al · 2024
Cited alongside, same era.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al · 2024
Cited alongside, same era.
Disentangling memory and reasoning ability in large language models
Jin, M., Luo, W., Cheng, S., Wang, X., Hua, W., Tang, R., Wang, W. Y., and Zhang, Y · 2024
Cited alongside, same era.
Training language models to self-correct via reinforcement learning
Kumar, A., Zhuang, V., Agarwal, R., Su, Y., Co-Reyes, J. D., Singh, A., Baumli, K., Iqbal, S., Bishop, C., Roelofs, R., et al · 2024
Cited alongside, same era.
Chen, H., Tu, H., Wang, F., Liu, H., Tang, X., Du, X., Zhou, Y., and Xie, C · 2025
Closest in time.
Sft memorizes, rl generalizes: A comparative study of foundation model post-training
Chu, T., Zhai, Y., Yang, J., Tong, S., Xie, S., Schuurmans, D., Le, Q. V., Levine, S., and Ma, Y · 2025
Closest in time.
Deng, Y., Bansal, H., Yin, F., Peng, N., Wang, W., and Chang, K.-W · 2025
Closest in time.
Medgemma hugging face
Google · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
m1: Unleash the potential of test-time scaling for medical reasoning with large language models
Huang, X., Wu, J., Liu, H., Tang, X., and Zhou, Y · 2025
Closest in time.
Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., Zhang, C. B. C., Shaaban, M., Ling, J., Shi, S., et al · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., et al · 2025
Closest in time.
Medreason: Eliciting factual medical reasoning steps in llms via knowledge graphs
Wu, J., Deng, W., Li, X., Liu, S., Mi, T., Peng, Y., Xu, Z., Liu, Y., Cho, H., Choi, C.-I., et al · 2025
Closest in time.
R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Yang, Y., He, X., Pan, H., Jiang, X., Deng, Y., Yang, X., Lu, H., Yin, D., Rao, F., Zhu, M., et al · 2025
Closest in time.
Easyr1: An efficient, scalable, multi-modality rl training framework, 2025
Zheng, Y., Lu, J., Wang, S., Feng, Z., Kuang, D., and Xiong, Y · 2025
Closest in time.
Medxpertqa: Benchmarking expert-level medical reasoning and understanding
Zuo, Y., Qu, S., Li, Y., Chen, Z., Zhu, X., Hua, E., Zhang, K., Ding, N., and Zhou, B · 2025
Closest in time.