Fetching the paper…
Reading the bibliography…
Evaluating generated radiology reports is crucial for the development of radiology AI, but existing metrics fail to reflect the task's clinical requirements.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Radpeer peer review: relevance, use, concerns, challenges, and direction forward
H. Abujudeh, R. S. Pyatt Jr, M. A. Bruno, A. L. Chetlen, D. Buck, S. K. Hobbs, C. Roth, C. Truwit, R. Agarwal, S. T. Kennedy, et al · 2014
Earlier work this paper cites.
A systematic review of the diagnostic accuracy of artificial intelligence-based computer programs to analyze chest x-rays for pulmonary tuberculosis
M. Harris, A. Qi, L. Jeagal, N. Torabi, D. Menzies, A. Korobitsyn, M. Pai, R. R. Nathavitharana, and F. Ahmad Khan · 2019
Earlier work this paper cites.
Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports
A. E. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P. Lungren, C.-y. Deng, R. G. Mark, and S. Horng · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Earlier work this paper cites.
Combining automatic labelers and expert annotations for accurate radiology report labeling using bert
A. Smit, S. Jain, P. Rajpurkar, A. Pareek, A. Y. Ng, and M. Lungren · 2020
Earlier work this paper cites.
Artificial intelligence solutions for analysis of x-ray images
S. J. Adams, R. D. E. Henderson, X. Yi, and P. Babyn · 2021
Earlier work this paper cites.
Radgraph: Extracting clinical entities and relations from radiology reports
S. Jain, A. Agrawal, A. Saporta, S. Truong, T. Bui, P. Chambon, Y. Zhang, M. P. Lungren, A. Y. Ng, C. Langlotz, et al · 2021
Earlier work this paper cites.
Should we replace radiologists with deep learning? pigeons, error and trust in medical ai
R. Alvarado · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
From sparse to dense: Gpt-4 summarization with chain of density prompting
G. Adams, A. Fabbri, F. Ladhak, E. Lehman, and N. Elhadad · 2023
Cited alongside, same era.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Cited alongside, same era.
Radiology-aware model-based evaluation metric for report generation
G-eval: Nlg evaluation using gpt-4 with better human alignment
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
P. Törnberg · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Chatcad: Interactive computer-aided diagnosis on medical image using large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Calamida, F. Nooralahzadeh, M. Rohanian, K. Fujimoto, M. Nishio, and M. Krauthammer · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
C.-H. Chiang and H.-y. Lee · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Cited alongside, same era.
Chatgpt outperforms crowd workers for text-annotation tasks
F. Gilardi, M. Alizadeh, and M. Kubli · 2023
Cited alongside, same era.
Maira-1: A specialised large multimodal model for radiology report generation
S. L. Hyland, S. Bannur, K. Bouzid, D. C. Castro, M. Ranjit, A. Schwaighofer, F. Pérez-García, V. Salvatelli, S. Srivastav, A. Thieme, et al · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
Radgraph2: Modeling disease progression in radiology reports via hierarchical information extraction
S. Khanna, A. Dejl, K. Yoon, S. Q. Truong, H. Duong, A. Saenz, and P. Rajpurkar · 2023
Cited alongside, same era.
S. Wang, Z. Zhao, X. Ouyang, Q. Wang, and D. Shen · 2023
Later among the works it cites.
Evaluating progress in automatic chest x-ray radiology report generation
F. Yu, M. Endo, R. Krishnan, I. Pan, A. Tsai, E. P. Reis, E. K. U. N. Fonseca, H. M. H. Lee, Z. S. H. Abad, A. Y. Ng, et al · 2023
Later among the works it cites.
Universalner: Targeted distillation from large language models for open named entity recognition
W. Zhou, S. Zhang, Y. Gu, M. Chen, and H. Poon · 2023
Later among the works it cites.
Openmedlm: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
A. Garikipati, J. Maharjan, N. P. Singh, L. Cyrus, M. Sharma, M. Ciobanu, G. Barnes, Q. Mao, and R. Das · 2024
Closest in time.
Large language model meets graph neural network in knowledge distillation
S. Hu, G. Zou, S. Yang, B. Zhang, and Y. Chen · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.