Fetching the paper…
Reading the bibliography…
LLMs have achieved significant performance progress in various NLP applications.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. Zhu · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
C. Lin and E. H. Hovy · 2003
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Q. Jin, B. Dhingra, Z. Liu, W. W. Cohen, and X. Lu · 2019
Earlier work this paper cites.
A hierarchical attention retrieval model for healthcare question answering
M. Zhu, A. Ahuja, W. Wei, and C. K. Reddy · 2019
Earlier work this paper cites.
D. Jin, E. Pan, N. Oufattole, W. Weng, H. Fang, and P. Szolovits · 2020
Earlier work this paper cites.
D. Jin, E. Pan, N. Oufattole, W. Weng, H. Fang, and P. Szolovits · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
All NLP tasks are generation tasks: A general pretraining framework
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Basic standards for medical record writing (trial)
NHC · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
A. Pal, L. K. Umapathi, and M. Sankarasubbu · 2022
Earlier work this paper cites.
Large language models encode clinical knowledge
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. K. Tanwani, H. Cole-Lewis, S. Pfohl, P. Payne, M. Seneviratne, P. Gamble, C. Kelly, N. Schärli, A. Chowdhery, P. A. Mansfield, B. A. y Arcas, D. R. Webster, G. S. Corrado, Y. Matias, K. Chou, J. Gottweis, N. Tomasev, Y. Liu, A. Rajkomar, J. K. Barral, C. Semturs, A. Karthikesalingam, and V. Natarajan · 2022
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y. Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y. Zhang, Z. Zhang, C. Zhou, J. Zhou, X. Zhou, and T. Zhu · 2023
Earlier work this paper cites.
Disc-medllm: Bridging general large language models and real-world medical consultation
Z. Bao, W. Chen, S. Xiao, K. Ren, J. Wu, C. Zhong, J. Peng, X. Huang, and Z. Wei · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. M. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang · 2023
Cited alongside, same era.
Huatuogpt-ii, one-stage training for medical adaption of llms
J. Chen, X. Wang, A. Gao, F. Jiang, S. Chen, H. Zhang, D. Song, W. Xie, C. Kong, J. Li, X. Wan, H. Li, and B. Wang · 2023
Cited alongside, same era.
Y. Chen, Z. Wang, X. Xing, H. Zheng, Z. Xu, K. Fang, J. Wang, S. Li, J. Wu, Q. Liu, and X. Xu · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini · 2023
Cited alongside, same era.
On large language models’ selection bias in multi-choice questions
C. Zheng, H. Zhou, F. Meng, J. Zhou, and M. Huang · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
A survey of large language models in medicine: Progress, application, and challenge
H. Zhou, B. Gu, X. Zou, Y. Li, S. S. Chen, P. Zhou, J. Liu, Y. Hua, C. Mao, X. Wu, Z. Li, and F. Liu · 2023
Later among the works it cites.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic · 2024
Closest in time.
Medbench: A large-scale chinese benchmark for evaluating medical large language models
Y. Cai, L. Wang, Y. Wang, G. de Melo, Y. Zhang, Y. Wang, and L. He · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benchmarking large language models on cmexam - A comprehensive chinese medical exam dataset
J. Liu, P. Zhou, Y. Hua, D. Chong, Z. Tian, A. Liu, H. Wang, C. You, Z. Guo, L. Zhu, and M. L. Li · 2023
Cited alongside, same era.
Taiyi: A bilingual fine-tuned large language model for diverse biomedical tasks
L. Luo, J. Ning, Y. Zhao, Z. Wang, Z. Ding, P. Chen, W. Fu, Q. Han, G. Xu, Y. Qiu, D. Pan, J. Li, H. Li, W. Feng, S. Tu, Y. Liu, Z. Yang, J. Wang, Y. Sun, and H. Lin · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Towards expert-level medical question answering with large language models
K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, L. Hou, K. Clark, S. Pfohl, H. Cole-Lewis, D. Neal, M. Schaekermann, A. Wang, M. Amin, S. Lachgar, P. A. Mansfield, S. Prakash, B. Green, E. Dominowska, B. A. y Arcas, N. Tomasev, Y. Liu, R. Wong, C. Semturs, S. S. Mahdavi, J. K. Barral, D. R. Webster, G. S. Corrado, Y. Matias, S. Azizi, A. Karthikesalingam, and V. Natarajan · 2023
Cited alongside, same era.
Medagents: Large language models as collaborators for zero-shot medical reasoning
X. Tang, A. Zou, Z. Zhang, Y. Zhao, X. Zhang, A. Cohan, and M. Gerstein · 2023
Cited alongside, same era.
Bluelm: An open multilingual 7b language model
B. Team · 2023
Cited alongside, same era.
CMB: A comprehensive medical benchmark in chinese
X. Wang, G. H. Chen, D. Song, Z. Zhang, Z. Chen, Q. Xiao, F. Jiang, J. Li, X. Wan, B. Wang, and H. Li · 2023
Cited alongside, same era.
Epidemic modeling with generative agents
R. Williams, N. Hosseinichimeh, A. Majumdar, and N. Ghaffarzadegan · 2023
Cited alongside, same era.
Z. Cai, M. Cao, H. Chen, K. Chen, K. Chen, X. Chen, X. Chen, Z. Chen, Z. Chen, P. Chu, X. Dong, H. Duan, Q. Fan, Z. Fei, Y. Gao, J. Ge, C. Gu, Y. Gu, T. Gui, A. Guo, Q. Guo, C. He, Y. Hu, T. Huang, T. Jiang, P. Jiao, Z. Jin, Z. Lei, J. Li, J. Li, L. Li, S. Li, W. Li, Y. Li, H. Liu, J. Liu, J. Hong, K. Liu, K. Liu, X. Liu, C. Lv, H. Lv, K. Lv, L. Ma, R. Ma, Z. Ma, W. Ning, L. Ouyang, J. Qiu, Y. Qu, F. Shang, Y. Shao, D. Song, Z. Song, Z. Sui, P. Sun, Y. Sun, H. Tang, B. Wang, G. Wang, J. Wang, J. Wang, R. Wang, Y. Wang, Z. Wang, X. Wei, Q. Weng, F. Wu, Y. Xiong, and et al · 2024
Closest in time.
Spark-3 website
L. Iflytek Co · 2024
Closest in time.
Adaptive collaboration strategy for llms in medical decision making
Y. Kim, C. Park, H. Jeong, Y. S. Chan, X. Xu, D. McDuff, C. Breazeal, and H. W. Park · 2024
Closest in time.
Agent hospital: A simulacrum of hospital with evolvable medical agents
J. Li, S. Wang, M. Zhang, W. Li, Y. Lai, X. Kang, W. Ma, and Y. Liu · 2024
Closest in time.
Towards building multilingual language model for medicine
P. Qiu, C. Wu, X. Zhang, W. Lin, H. Wang, Y. Zhang, Y. Wang, and W. Xie · 2024
Closest in time.
Wingpt2 website
W. H. A. Research · 2024
Closest in time.
Towards large language models as copilots for theorem proving in lean
P. Song, K. Yang, and A. Anandkumar · 2024
Closest in time.
Clinical text datasets for medical artificial intelligence and large language models—a systematic review
J. Wu, X. Liu, M. Li, W. Li, Z. Su, S. Lin, L. Garay, Z. Zhang, Y. Zhang, Q. Zeng, et al · 2024
Closest in time.
Yi: Open foundation models by 01.ai
A. Young, B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, H. Li, J. Zhu, J. Chen, J. Chang, K. Yu, P. Liu, Q. Liu, S. Yue, S. Yang, S. Yang, T. Yu, W. Xie, W. Huang, X. Hu, X. Ren, X. Niu, P. Nie, Y. Xu, Y. Liu, Y. Wang, Y. Cai, Z. Gu, Z. Liu, and Z. Dai · 2024
Closest in time.