Fetching the paper…
Reading the bibliography…
In the rapidly evolving landscape of large language models (LLMs) for medical applications, ensuring the reliability and accuracy of these models in clinical settings is paramount.
BERTScore: Evaluating Text Generation with BERT
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2019 · 1904
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Zhao, W.; Peyrard, M.; Liu, F.; Gao, Y.; Meyer, C. M.; and Eger, S. 2019 · 1909
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Wang, A.; Cho, K.; and Lewis, M. 2020 · 2004
Earlier work this paper cites.
USR: An unsupervised and reference free evaluation metric for dialog generation
Mehri, S.; and Eskenazi, M. 2020 · 2005
Earlier work this paper cites.
Pearson correlation coefficient
Cohen, I.; Huang, Y.; Chen, J.; Benesty, J.; Benesty, J.; Chen, J.; Huang, Y.; and Cohen, I. 2009 · 2009
Earlier work this paper cites.
GRADE: Automatic graph-enhanced coherence metric for evaluating open-domain dialogue systems
Huang, L.; Ye, Z.; Qin, J.; Lin, L.; and Liang, X. 2020 · 2010
Earlier work this paper cites.
A guideline of selecting and reporting intraclass correlation coefficients for reliability research
Koo, T. K.; and Li, M. Y. 2016 · 2016
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
MedDialog: Large-scale medical dialogue datasets
Zeng, G.; Yang, W.; Ju, Z.; Yang, Y.; Wang, S.; Zhang, R.; Zhou, M.; Zeng, J.; Dong, X.; Zhang, R.; et al. 2020 · 2020
Cited alongside, same era.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T.; Fu, D.; Ermon, S.; Rudra, A.; and Ré, C. 2022 · 2022
Cited alongside, same era.
Results of WMT22 metrics shared task: Stop using BLEU–neural metrics are better and more robust
Freitag, M.; Rei, R.; Mathur, N.; Lo, C.-k.; Stewart, C.; Avramidis, E.; Kocmi, T.; Foster, G.; Lavie, A.; and Martins, A. F. 2022 · 2022
Cited alongside, same era.
MedMCQA: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A.; Umapathi, L. K.; and Sankarasubbu, M. 2022 · 2022
Cited alongside, same era.
Branch-solve-merge improves large language model evaluation and generation
Saha, S.; Levy, O.; Celikyilmaz, A.; Bansal, M.; Weston, J.; and Li, X. 2023 · 2023
Later among the works it cites.
Style over substance: Evaluation biases for large language models
Wu, M.; and Aji, A. F. 2023 · 2023
Later among the works it cites.
Towards Effective Automatic Evaluation of Generated Reflections for Motivational Interviewing
Wu, Z.; Helaoui, R.; Reforgiato Recupero, D.; and Riboni, D. 2023 · 2023
Later among the works it cites.
PyTorch FSDP: Experiences on scaling fully sharded data parallel
Zhao, Y.; Gu, A.; Varma, R.; Luo, L.; Huang, C.-C.; Xu, M.; Wright, L.; Shojanazeri, H.; Ott, M.; Shleifer, S.; et al. 2023 · 2023
Later among the works it cites.
MedBench: A large-scale Chinese benchmark for evaluating medical large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards a unified multi-dimensional evaluator for text generation
Zhong, M.; Liu, Y.; Yin, D.; Mao, Y.; Jiao, Y.; Liu, P.; Zhu, C.; Ji, H.; and Han, J. 2022 · 2022
Cited alongside, same era.
Chen, Y.; Wang, R.; Jiang, H.; Shi, S.; and Xu, R. 2023 · 2023
Cited alongside, same era.
Large language models are state-of-the-art evaluators of translation quality
Kocmi, T.; and Federmann, C. 2023 · 2023
Cited alongside, same era.
Capabilities of GPT-4 on medical challenge problems
Nori, H.; King, N.; McKinney, S. M.; Carignan, D.; and Horvitz, E. 2023 · 2023
Cited alongside, same era.
Is ChatGPT a good NLG evaluator? A preliminary study
Wang, J.; Liang, Y.; Meng, F.; Sun, Z.; Shi, H.; Li, Z.; Xu, J.; Qu, J.; and Zhou, J. 2023a
Cited in the paper.
Large language models are not fair evaluators
Wang, P.; Li, L.; Chen, L.; Zhu, D.; Lin, B.; Cao, Y.; Liu, Q.; Liu, T.; and Sui, Z. 2023b
Cited in the paper.
PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization
Wang, Y.; Yu, Z.; Zeng, Z.; Yang, L.; Wang, C.; Chen, H.; Jiang, C.; Xie, R.; Wang, J.; Xie, X.; et al. 2023c
Cited in the paper.
Cai, Y.; Wang, L.; Wang, Y.; de Melo, G.; Zhang, Y.; Wang, Y.; and He, L. 2024 · 2024
Later among the works it cites.
GPTScore: Evaluate as You Desire
Fu, J.; Ng, S.-K.; Jiang, Z.; and Liu, P. 2024 · 2024
Later among the works it cites.
Large language models are inconsistent and biased evaluators
Stureborg, R.; Alikaniotis, D.; and Suhara, Y. 2024 · 2024
Later among the works it cites.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2024 · 2024
Later among the works it cites.