Fetching the paper…
Reading the bibliography…
Surgery requires comprehensive medical knowledge, visual assessment skills, and procedural expertise.
You only look once: Unified, real-time object detection
Redmon, J., Divvala, S., Girshick, R. & Farhadi, A · 2016
Earlier work this paper cites.
Yolo9000: better, faster, stronger
Redmon, J. & Farhadi, A · 2017
Earlier work this paper cites.
Artificial intelligence in surgery: promises and perils
Hashimoto, D. A., Rosman, G., Rus, D. & Meireles, O. R · 2018
Earlier work this paper cites.
Unified vision-language pre-training for image captioning and vqa
Zhou, L. et al · 2020
Earlier work this paper cites.
Yolov4: Optimal speed and accuracy of object detection
Bochkovskiy, A., Wang, C.-Y. & Liao, H.-Y. M · 2020
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D. et al · 2021
Earlier work this paper cites.
Scaling up vision-language pre-training for image captioning
Hu, X. et al · 2022
Earlier work this paper cites.
Roentgen: vision-language foundation model for chest x-ray generation
Chambon, P. et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B. et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L. et al · 2022
Earlier work this paper cites.
Surgical-vqa: Visual question answering in surgical scenes using transformer
Seenivasan, L., Islam, M., Krishna, A. K. & Ren, H · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., Umapathi, L. K. & Sankarasubbu, M · 2022
Earlier work this paper cites.
Data splits and metrics for method benchmarking on surgical action triplet datasets
Nwoye, C. I. & Padoy, N · 2022
Earlier work this paper cites.
Cholectriplet2021: A benchmark challenge for surgical action triplet recognition
Nwoye, C. I. et al · 2022
Earlier work this paper cites.
Artificial intelligence in surgical learning
Pakkasjärvi, N., Luthra, T. & Anand, S · 2023
Earlier work this paper cites.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Yang, A. et al · 2023
Earlier work this paper cites.
Prompting large language models with answer heuristics for knowledge-based visual question answering
Shao, Z., Yu, Z., Wang, M. & Yu, J · 2023
Earlier work this paper cites.
From images to textual prompts: Zero-shot visual question answering with frozen large language models
Guo, J. et al · 2023
Earlier work this paper cites.
Image captioning for effective use of language models in knowledge-based visual question answering
Salaberria, A., Azkune, G., de Lacalle, O. L., Soroa, A. & Agirre, E · 2023
Earlier work this paper cites.
A foundational multimodal vision language ai assistant for human pathology
Lu, M. Y. et al · 2023
Earlier work this paper cites.
A visual–language foundation model for pathology image analysis using medical twitter
Huang, Z., Bianchi, F., Yuksekgonul, M., Montine, T. J. & Zou, J · 2023
Cited alongside, same era.
Pellegrini, C., Özsoy, E., Busam, B., Navab, N. & Keicher, M · 2023
Cited alongside, same era.
Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery
Seenivasan, L., Islam, M., Kannan, G. & Ren, H · 2023
Cited alongside, same era.
Surgical-vqla: Transformer with gated vision-language embedding for visual question localized-answering in robotic surgery
Bai, L., Islam, M., Seenivasan, L. & Ren, H · 2023
Cited alongside, same era.
Segment anything
Kirillov, A. et al · 2023
Cited alongside, same era.
General surgery vision transformer: A video pre-trained foundation model for general surgery
Schmidgall, S., Kim, J. W., Jopling, J. & Krieger, A · 2024
Closest in time.
Robots learning to imitate surgeons—challenges and possibilities
Schmidgall, S., Kim, J. W. & Krieger, A · 2024
Closest in time.
Visual instruction tuning
Liu, H., Li, C., Wu, Q. & Lee, Y. J · 2024
Closest in time.
Improved baselines with visual instruction tuning
Liu, H., Li, C., Li, Y. & Lee, Y. J · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge (2024)
Liu, H. et al · 2024
Closest in time.
Vision-language models for vision tasks: A survey
Zhang, J., Huang, J., Jin, S. & Lu, S · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Segment anything model for medical image analysis: an experimental study
Mazurowski, M. A. et al · 2023
Cited alongside, same era.
Instruction tuning for large language models: A survey
Zhang, S. et al · 2023
Cited alongside, same era.
Med-flamingo: a multimodal medical few-shot learner
Moor, M. et al · 2023
Cited alongside, same era.
Can generalist foundation models outcompete special-purpose tuning? case study in medicine
Nori, H. et al · 2023
Cited alongside, same era.
Language models are susceptible to incorrect patient self-diagnosis in medical applications
Ziaei, R. & Schmidgall, S · 2023
Cited alongside, same era.
Capabilities of gpt-4 on medical challenge problems
Nori, H., King, N., McKinney, S. M., Carignan, D. & Horvitz, E · 2023
Cited alongside, same era.
Achiam, J. et al · 2023
Cited alongside, same era.
A visual-language foundation model for computational pathology
Lu, M. Y. et al · 2024
Closest in time.
Wang, G. et al · 2024
Closest in time.
Segment anything in high quality
Ke, L. et al · 2024
Closest in time.
Scaling instruction-finetuned language models
Chung, H. W. et al · 2024
Closest in time.
Capabilities of gemini models in medicine
Saab, K. et al · 2024
Closest in time.
Almanac—retrieval-augmented language models for clinical medicine
Zakka, C. et al · 2024
Closest in time.
Addressing cognitive bias in medical language models
Schmidgall, S. et al · 2024
Closest in time.
Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments
Schmidgall, S. et al · 2024
Closest in time.
Sufia: Language-guided augmented dexterity for robotic surgical assistants
Moghani, M. et al · 2024
Closest in time.
Advancing surgical vqa with scene graph knowledge
Yuan, K. et al · 2024
Closest in time.
Surgical robot transformer (srt): Imitation learning for surgical tasks
Kim, J. W. et al · 2024
Closest in time.
Orbit-surgical: An open-simulation framework for learning surgical augmented dexterity
Yu, Q. et al · 2024
Closest in time.
Openvla: An open-source vision-language-action model
Kim, M. J. et al · 2024
Closest in time.
A survey on vision-language-action models for embodied ai
Ma, Y., Song, Z., Zhuang, Y., Hao, J. & King, I · 2024
Closest in time.