Fetching the paper…
Reading the bibliography…
Generalist foundation models (GFMs) are renowned for their exceptional capability and flexibility in effectively generalizing across diverse tasks and modalities.
Jacobs RA, Jordan MI, Nowlan SJ, et al (1991) Adaptive mixtures of local experts. Neural computation 3(1):79–87
1991
Earlier work this paper cites.
Banerjee S, Lavie A (2005) Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp 65–72
2005
Earlier work this paper cites.
Deng J, Dong W, Socher R, et al (2009) Imagenet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Ieee, pp 248–255
2009
Earlier work this paper cites.
Demner-Fushman D, Antani S, Simpson M, et al (2012) Design and development of a multimodal biomedical information retrieval system. Journal of Computing Science and Engineering 6(2):168–177
2012
Earlier work this paper cites.
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25
2012
Earlier work this paper cites.
Nguyen HQ, Lam K, Le LT, et al (2020) Vindr-cxr: An open dataset of chest x-rays with radiologist’s annotations. 2012.15029
2012
Earlier work this paper cites.
Lin TY, Maire M, Belongie S, et al (2014) Microsoft coco: Common objects in context. In: Proceedings of European Conference on Computer Vision, Springer, pp 740–755
2014
Earlier work this paper cites.
Silva J, Histace A, Romain O, et al (2014) Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer. International journal of computer assisted radiology and surgery 9:283–293
2014
Earlier work this paper cites.
Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:14091556
2014
Earlier work this paper cites.
Antol S, Agrawal A, Lu J, et al (2015) Vqa: Visual question answering. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
Earlier work this paper cites.
Demner-Fushman D, Kohli MD, Rosenman MB, et al (2016) Preparing a collection of radiology examinations for distribution and retrieval. Journal of the American Medical Informatics Association 23(2):304–310
2016
Earlier work this paper cites.
He K, Zhang X, Ren S, et al (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 770–778
2016
Earlier work this paper cites.
Sawyer-Lee R, Gimenez F, Hoogi A, et al (2016) Curated breast imaging subset of digital database for screening mammography (cbis-ddsm). The Cancer Imaging Archive, 10.7937/K9/TCIA.2016.7O02S9CY
2016
Earlier work this paper cites.
Bejnordi BE, Veta M, Van Diest PJ, et al (2017) Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. Jama 318(22):2199–2210
2017
Earlier work this paper cites.
Huang G, Liu Z, Van Der Maaten L, et al (2017) Densely connected convolutional networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 4700–4708
2017
Earlier work this paper cites.
Pogorelov K, Randel KR, Griwodz C, et al (2017) Kvasir: A multi-class image dataset for computer aided gastrointestinal disease detection. In: Proceedings of the 8th ACM on Multimedia Systems Conference, pp 164–169
2017
Earlier work this paper cites.
Kawahara J, Daneshvar S, Argenziano G, et al (2018) Seven-point checklist and skin lesion classification using multitask multimodal neural nets. IEEE journal of biomedical and health informatics 23(2):538–546
2018
Earlier work this paper cites.
Lau JJ, Gayen S, Demner D, et al (2018) Visual question answering in radiology (vqa-rad). Open Science Framework
2018
Earlier work this paper cites.
Tschandl P, Rosendahl C, Kittler H (2018) The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific data 5(1):1–9
2018
Earlier work this paper cites.
Ben Abacha A, Hasan SA, Datla VV, et al (2019) Vqa-med: Overview of the medical visual question answering task at imageclef 2019. In: Working Notes of CLEF 2019, CEUR Workshop Proceedings, vol 2380. CEUR-WS.org, Lugano, Switzerland, URL https://ceur-ws.org/Vol-2380/paper_272.pdf
2019
Earlier work this paper cites.
Irvin J, Rajpurkar P, Ko M, et al (2019) Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In: Proceedings of the AAAI conference on artificial intelligence, pp 590–597
2019
Earlier work this paper cites.
Johnson AE, Pollard TJ, Berkowitz SJ, et al (2019) Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data 6(1):317
2019
Earlier work this paper cites.
Tan M, Le Q (2019) Efficientnet: Rethinking model scaling for convolutional neural networks. In: Proceedings of the International Conference on Machine Learning, PMLR, pp 6105–6114
2019
Earlier work this paper cites.
Chen Z, Song Y, Chang TH, et al (2020) Generating radiology reports via memory-driven transformer. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing
2020
Earlier work this paper cites.
Goel S (2020) Dermnet. https://www.kaggle.com/datasets/shubhamgoel27/dermnet
2020
Earlier work this paper cites.
He X, Zhang Y, Mou L, et al (2020) Pathvqa: 30000+ questions for medical visual question answering. arXiv preprint arXiv:200310286
2020
Earlier work this paper cites.
Lewis P, Perez E, Piktus A, et al (2020) Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems 33:9459–9474
2020
Cited alongside, same era.
Li N, Li T, Hu C, et al (2021) A benchmark of ocular disease intelligent recognition: One shot for multi-disease detection. In: Benchmarking, Measuring, and Optimizing: Third BenchCouncil International Symposium, Bench 2020, Virtual Event, November 15–16, 2020, Revised Selected Papers 3, Springer, pp 177–193
2020
Cited alongside, same era.
Pacheco AG, Lima GR, Salomao AS, et al (2020) Pad-ufes-20: A skin lesion dataset composed of patient data and clinical images collected from smartphones. Data in brief 32:106221
2020
Cited alongside, same era.
Caron M, Touvron H, Misra I, et al (2021) Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp 9650–9660
2021
Cited alongside, same era.
Liu H, Li C, Wu Q, et al (2023) Visual instruction tuning. In: Advances in Neural Information Processing Systems, vol 36. Curran Associates, Inc., pp 34892–34916, URL https://proceedings.neurips.cc/paper_files/paper/2023/file/6dcf277ea32ce3288914faf369fe6de0-Paper-Conference.pdf
2023
Later among the works it cites.
Nakayama LF, Goncalves M, Zago Ribeiro L, et al (2023) A brazilian multilabel ophthalmological dataset (brset). PhysioNet https://doi org/10 13026
2023
Later among the works it cites.
Nguyen HT, Nguyen HQ, Pham HH, et al (2023) Vindr-mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography. Scientific Data 10(1):277
2023
Later among the works it cites.
Panchal S, Naik A, Kokare M, et al (2023) Retinal fundus multi-disease image dataset (rfmid) 2.0: A dataset of frequently and rarely identified diseases. Data 8(2). 10.3390/data8020029 , URL https://www.mdpi.com/2306-5729/8/2/29
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cen LP, Ji J, Lin JW, et al (2021) Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks. Nature communications 12(1):4828
2021
Cited alongside, same era.
Dosovitskiy A, Beyer L, Kolesnikov A, et al (2021) An image is worth 16x16 words: Transformers for image recognition at scale. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=YicbFdNTTy
2021
Cited alongside, same era.
Liu B, Zhan LM, Xu L, et al (2021) Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), IEEE, pp 1650–1654
2021
Cited alongside, same era.
Nguyen HT, Pham HH, Nguyen NT, et al (2021) Vindr-spinexr: A deep learning framework for spinal lesions detection and classification from radiographs. In: Proceedings of the Medical Image Computing and Computer Assisted Intervention (MICCAI), Springer, pp 291–301
2021
Cited alongside, same era.
Radford A, Kim JW, Hallacy C, et al (2021) Learning transferable visual models from natural language supervision. In: Proceedings of the International Conference on Machine Learning, PMLR, pp 8748–8763
2021
Cited alongside, same era.
Rajbhandari S, Ruwase O, Rasley J, et al (2021) Zero-infinity: breaking the gpu memory wall for extreme scale deep learning. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Association for Computing Machinery, 10.1145/3458817.3476205 , URL https://doi.org/10.1145/3458817.3476205
2021
Cited alongside, same era.
Rotemberg V, Kurtansky N, Betz-Stablein B, et al (2021) A patient-centric dataset of images and metadata for identifying melanomas using clinical context. Scientific data 8(1):34
2021
Cited alongside, same era.
Smedsrud PH, Thambawita V, Hicks SA, et al (2021) Kvasir-Capsule, a video capsule endoscopy dataset. Scientific Data 8(1):142. 10.1038/s41597-021-00920-z
2021
Cited alongside, same era.
Touvron H, Martin L, Stone K, et al (2023) Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:230709288
2023
Later among the works it cites.
Van Veen D, Van Uden C, Attias M, et al (2023) Radadapt: Radiology report summarization via lightweight domain adaptation of large language models. arXiv preprint arXiv:230501146
2023
Later among the works it cites.
Wang D, Wang X, Wang L, et al (2023) A real-world dataset and benchmark for foundation model adaptation in medical image classification. Scientific Data 10(1):574
2023
Later among the works it cites.
Yang J, Shi R, Wei D, et al (2023) Medmnist v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification. Scientific Data 10(1):41
2023
Later among the works it cites.
Yu F, Endo M, Krishnan R, et al (2023) Evaluating progress in automatic chest x-ray radiology report generation. Patterns 4(9)
2023
Later among the works it cites.
Zhu D, Chen J, Shen X, et al (2023) Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:230410592
2023
Later among the works it cites.
Doerrich S, Di Salvo F, Brockmann J, et al (2024) Rethinking model prototyping through the medmnist+ dataset collection. arXiv preprint arXiv:240415786
2024
Closest in time.
Hu Y, Li T, Lu Q, et al (2024) Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm. arXiv preprint arXiv:240209181
2024
Closest in time.
Jin H, Che H, Lin Y, et al (2024) Promptmrg: Diagnosis-driven prompts for medical report generation. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp 2607–2615
2024
Closest in time.
Li C, Wong C, Zhang S, et al (2024) Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Luo L, Wu M, Li M, et al (2024) Towards non-invasive and personalized management of breast cancer patients from multiparametric mri via a large mixture-of-modality-experts model. arXiv preprint arXiv:240812606
2024
Closest in time.
Royer C, Menze B, Sekuboyina A (2024) Multimedeval: A benchmark and a toolkit for evaluating medical vision-language models. 2402.09262
2024
Closest in time.
Saab K, Tu T, Weng WH, et al (2024) Capabilities of gemini models in medicine. arXiv preprint arXiv:240418416
2024
Closest in time.
Tu T, Azizi S, Driess D, et al (2024) Towards generalist biomedical ai. NEJM AI 1(3):AIoa2300138
2024
Closest in time.
Tung C, Lin Y, Yin J, et al (2024) Exploring vision language pretraining with knowledge enhancement via large language model. In: International Workshop on Trustworthy Artificial Intelligence for Healthcare, Springer, pp 81–91
2024
Closest in time.
Xiong C, Chen H, Zheng H, et al (2024) Mome: Mixture of multimodal experts for cancer survival prediction. arXiv preprint arXiv:240609696
2024
Closest in time.
Xu Y, Wang Y, Zhou F, et al (2024) A multimodal knowledge-enhanced whole-slide pathology foundation model. arXiv preprint arXiv:240715362
2024
Closest in time.
Yang L, Xu S, Sellergren A, et al (2024) Advancing multimodal medical capabilities of gemini. arXiv preprint arXiv:240503162
2024
Closest in time.
Zhang K, Zhou R, Adhikarla E, et al (2024) A generalist vision–language foundation model for diverse biomedical tasks. Nature Medicine pp 1–13
2024
Closest in time.
Zhou HY, Adithan S, Acosta JN, et al (2024) A generalist learner for multifaceted medical image interpretation. arXiv preprint arXiv:240507988
2024
Closest in time.
Wang X, Peng Y, Lu L, et al (2017) Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 2097–2106
2097
Closest in time.