Fetching the paper…
Reading the bibliography…
In this work, we introduce MedAgentSim, an open-source simulated clinical environment with doctor, patient, and measurement agents designed to evaluate and enhance LLM performance in dynamic diagnostic settings.
Phaser Studio, I.: A fast, fun and free open source html5 game framework (2018), https://phaser.io/
2018
Earlier work this paper cites.
Lindeijer, T.: A free and open source, easy to use, and flexible full-featured level editor. (2019), https://www.mapeditor.org/
2019
Earlier work this paper cites.
Jin, D., Pan, E., Oufattole, N., Weng, W.H., Fang, H., Szolovits, P.: What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences 11
2021
Earlier work this paper cites.
Meyer, A.N., Giardina, T.D., Khawaja, L., Singh, H.: Patient and clinician experiences of uncertainty in the diagnostic process: current understanding and future directions. Patient Education and Counseling 104
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, A., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning (2021), available at https://openai.com/research/clip
2021
Earlier work this paper cites.
OpenAI: Chatgpt-3.5 (2022), available at https://openai.com/blog/chatgpt-3-5
2022
Earlier work this paper cites.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al.: Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35
2022
Earlier work this paper cites.
Zhong, C., Liao, K., Chen, W., Liu, Q., Peng, B., Huang, X., Peng, J., Wei, Z.: Hierarchical reinforcement learning for automatic disease diagnosis. Bioinformatics 38
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Johnson, A.E., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T.J., Hao, S., Moody, B., Gow, B., et al.: Mimic-iv, a freely accessible electronic health record dataset. Scientific data 10
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural information processing systems 36
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
OpenAI: Chatgpt-4 (2023), available at https://openai.com/blog/chatgpt-4
2023
Cited alongside, same era.
Park, J.S., O’Brien, J., Cai, C.J., Morris, M.R., Liang, P., Bernstein, M.S.: Generative agents: Interactive simulacra of human behavior. In: Proceedings of the 36th annual acm symposium on user interface software and technology. pp. 1–22 (2023)
2023
Cited alongside, same era.
Singhal, K., Azizi, S., Tu, T., Mahdavi, S.S., Wei, J., Chung, H.W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al.: Large language models encode clinical knowledge. Nature 620
2023
Cited alongside, same era.
2024
Later among the works it cites.
Liévin, V., Hother, C.E., Motzfeldt, A.G., Winther, O.: Can large language models reason about medical questions? Patterns 5
2024
Later among the works it cites.
Meta: Llama 3.3-70b instruct (2024), available at https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct
2024
Later among the works it cites.
OpenAI: Chatgpt-4o (2024), available at https://openai.com/blog/chatgpt-4o
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Vaid, A., Landi, I., Nadkarni, G., Nabeel, I.: Using fine-tuned large language models to parse clinical notes in musculoskeletal pain disorders. The Lancet Digital Health 5
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Anthropic: Introducing claude 3.5 sonnet (2024), available at https://www.anthropic.com/news/claude-3-5-sonnet
2024
Cited alongside, same era.
Community, H.F.: Huggingface (2024), https://huggingface.co/models
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
Team, Q.: Qwen2.5-72b: A 72 billion parameter language model (2024), available at https://huggingface.co/Qwen/Qwen2.5-72B
2024
Later among the works it cites.
2024
Later among the works it cites.
Wu, C., Lin, W., Zhang, X., Zhang, Y., Xie, W., Wang, Y.: Pmc-llama: toward building open-source language models for medicine. Journal of the American Medical Informatics Association p. ocae045 (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
Fan, Z., Wei, L., Tang, J., Chen, W., Siyuan, W., Wei, Z., Huang, F.: Ai hospital: Benchmarking large language models in a multi-agent medical interaction simulator. In: Proceedings of the 31st International Conference on Computational Linguistics. pp. 10183–10213 (2025)
2025
Closest in time.
Team, M.A.: Mistral small 3: A latency-optimized 24b-parameter model (2025), available at https://mistral.ai/news/mistral-small-3
2025
Closest in time.