Fetching the paper…
Reading the bibliography…
We introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning.
Differential Diagnosis of Common Complaints E-Book
Seller, R. H. and Symons, A. B · 2011
Earlier work this paper cites.
Human Anatomy and Physiology Preparatory Course
Liachovitzky, C · 2015
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Lau, J. J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D · 2018
Earlier work this paper cites.
Vqa-med: Overview of the medical visual question answering task at imageclef 2019
Ben Abacha, A., Hasan, S. A., Datla, V. V., Demner-Fushman, D., and Müller, H · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X · 2019
Earlier work this paper cites.
Reasoning processes in clinical reasoning: from the perspective of cognitive psychology
Shin, H. S · 2019
Earlier work this paper cites.
Five decades of research and theorization on clinical reasoning: a critical review
Yazdani, S. and Hoseini Abardeh, M · 2019
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing, 2020
Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., and Poon, H · 2020
Earlier work this paper cites.
Pathvqa: 30000+ questions for medical visual question answering
He, X., Zhang, Y., Mou, L., Xing, E., and Xie, P · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P · 2021
Earlier work this paper cites.
Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering
Liu, B., Zhan, L.-M., Xu, L., Ma, L., Yang, Y., and Wu, X.-M · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., Umapathi, L. K., and Sankarasubbu, M · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval
Jin, Q., Kim, W., Chen, Q., Comeau, D. C., Yeganova, L., Wilbur, W. J., and Lu, Z · 2023
Cited alongside, same era.
Pmc-vqa: Visual instruction tuning for medical visual question answering
Zhang, X., Wu, C., Zhao, Z., Lin, W., Zhang, Y., Wang, Y., and Xie, W · 2023
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Cited alongside, same era.
Internlm2 technical report, 2024
Cai, Z., Cao, M., Chen, H., Chen, K., Chen, K., Chen, X., Chen, X., Chen, Z., Chen, Z., Chu, P., Dong, X., Duan, H., Fan, Q., Fei, Z., Gao, Y., Ge, J., Gu, C., Gu, Y., Gui, T., Guo, A., Guo, Q., He, C., Hu, Y., Huang, T., Jiang, T., Jiao, P., Jin, Z., Lei, Z., Li, J., Li, J., Li, L., Li, S., Li, W., Li, Y., Liu, H., Liu, J., Hong, J., Liu, K., Liu, K., Liu, X., Lv, C., Lv, H., Lv, K., Ma, L., Ma, R., Ma, Z., Ning, W., Ouyang, L., Qiu, J., Qu, Y., Shang, F., Shao, Y., Song, D., Song, Z., Sui, Z., Sun, P., Sun, Y., Tang, H., Wang, B., Wang, G., Wang, J., Wang, J., Wang, R., Wang, Y., Wang, Z., Wei, X., Weng, Q., Wu, F., Xiong, Y., Xu, C., Xu, R., Yan, H., Yan, Y., Yang, X., Ye, H., Ying, H., Yu, J., Yu, J., Zang, Y., Zhang, C., Zhang, L., Zhang, P., Zhang, P., Zhang, R., Zhang, S., Zhang, S., Zhang, W., Zhang, W., Zhang, X., Zhang, X., Zhao, H., Zhao, Q., Zhao, X., Zhou, F., Zhou, Z., Zhuo, J., Zou, Y., Qiu, X., Qiao, Y., and Lin, D · 2024
Capabilities of gemini models in medicine
Saab, K., Tu, T., Weng, W.-H., Tanno, R., Stutz, D., Wulczyn, E., Zhang, F., Strother, T., Park, C., Vedadi, E., et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Later among the works it cites.
A comparative study on reasoning patterns of openai’s o1 model
Wu, S., Peng, Z., Du, X., Zheng, T., Liu, M., Wu, J., Ma, J., Li, Y., Yang, J., Zhou, W., et al · 2024
Later among the works it cites.
A preliminary study of o1 in medicine: Are we closer to an ai doctor?
Xie, Y., Wu, J., Tu, H., Yang, S., Zhao, B., Zong, Y., Jin, Q., Xie, C., and Zhou, Y · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Gemini 2.0 flash
Google · 2024
Cited alongside, same era.
Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm
Hu, Y., Li, T., Lu, Q., Shao, W., He, J., Qiao, Y., and Luo, P · 2024
Cited alongside, same era.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al · 2024
Cited alongside, same era.
Rho-1: Not all tokens are what you need
Lin, Z., Gou, Z., Gong, Y., Liu, X., Shen, Y., Xu, R., Lin, C., Yang, Y., Jiao, J., Duan, N., et al · 2024
Cited alongside, same era.
From medprompt to o1: Exploration of run-time strategies for medical challenge problems and beyond
Nori, H., Usuyama, N., King, N., McKinney, S. M., Fernandes, X., Zhang, S., and Horvitz, E · 2024
Cited alongside, same era.
Xu, R., Wang, Z., Fan, R.-Z., and Liu, P · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Ultramedical: Building specialized generalists in biomedicine
Zhang, K., Zeng, S., Hua, E., Ding, N., Chen, Z.-R., Ma, Z., Li, H., Cui, G., Qi, B., Zhu, X., et al · 2024
Later among the works it cites.
Zhu, K., Zheng, Y., and Chan, K. C. G · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
O1 replication journey–part 3: Inference-time scaling for medical reasoning
Huang, Z., Geng, G., Hua, S., Huang, Z., Zou, H., Zhang, S., Liu, P., and Zhang, X · 2025
Closest in time.
Gpt-o3-mini
OpenAI · 2025
Closest in time.
Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., Zhang, C. B. C., Shaaban, M., Ling, J., Shi, S., et al · 2025
Closest in time.
Qwen2.5-vl, January 2025
Team, Q · 2025
Closest in time.
Medreason: Eliciting factual medical reasoning steps in llms via knowledge graphs
Wu, J., Deng, W., Li, X., Liu, S., Mi, T., Peng, Y., Xu, Z., Liu, Y., Cho, H., Choi, C.-I., et al · 2025
Closest in time.