Fetching the paper…
Reading the bibliography…
Reasoning is a critical frontier for advancing medical image analysis, where transparency and trustworthiness play a central role in both clinician trust and regulatory approval.
Christiano, P.F., Leike, J., Brown, T., Martic, M., Legg, S., Amodei, D.: Deep reinforcement learning from human preferences. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al.: Mastering the game of go without human knowledge. nature 550
2017
Earlier work this paper cites.
Lau, J.J., Gayen, S., Ben Abacha, A., et al.: A dataset of clinically generated visual questions and answers about radiology images. Scientific data 5
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
Liu, B., Zhan, L.M., Xu, L., Ma, L., Yang, Y., Wu, X.M.: Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In: international symposium on biomedical imaging (ISBI). pp. 1650–1654 (2021)
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.: Training language models to follow instructions with human feedback. Advances in neural information processing systems 35
2022
Earlier work this paper cites.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al.: Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Akhter, Y., Singh, R., Vatsa, M.: Ai-based radiodiagnosis using chest x-rays: A review. Frontiers in big data 6
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., Gao, J.: Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
Qwen-Team: Qwq: Reflect deeply on the boundaries of the unknown (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Chen, J., Ouyang, R., Gao, A., Chen, S., Chen, G.H., Wang, X., Zhang, R., Cai, Z., Ji, K., Yu, G., Wan, X., Wang, B.: Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale (2024)
2024
Cited alongside, same era.
Hartsock, I., Rasool, G.: Vision-language models for medical report generation and visual question answering: A review. Frontiers in Artificial Intelligence 7
2024
Cited alongside, same era.
Hu, Y., Li, T., Lu, Q., Shao, W., He, J., Qiao, Y., Luo, P.: Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm. In: Conference on Computer Vision and Pattern Recognition. pp. 22170–22183 (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Zhang, K., Zhou, R., Adhikarla, E., Yan, Z., Liu, Y., Yu, J., Liu, Z., Chen, X., Davison, B.D., Ren, H., et al.: A generalist vision–language foundation model for diverse biomedical tasks. Nature Medicine pp. 1–13 (2024)
2024
Later among the works it cites.
Anthropic: Claude 3.7 sonnet system card (2025)
2025
Closest in time.
Chen, L., Li, L., Zhao, H., Song, Y., Vinci: R1-v: Reinforcing super generalization ability in vision-language models with less than $3. https://github.com/Deep-Agent/R1-V (2025), accessed: 2025-02-02
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
LMMs-Lab: open-r1-multimodal. https://github.com/EvolvingLMMs-Lab/open-r1-multimodal (2025), accessed: 2025-01-27
2025
Closest in time.
Shen, H., Zhang, Z., Zhang, Q., Xu, R., Zhao, T.: Vlm-r1: A stable and generalizable r1-style large vision-language model. https://github.com/om-ai-lab/VLM-R1 (2025), accessed: 2025-02-15
2025
Closest in time.