Fetching the paper…
Reading the bibliography…
Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al., 2020 · 1901
Earlier work this paper cites.
2017 robotic instrument segmentation challenge
Allan, M., Shvets, A., Kurmann, T., Zhang, Z., Duggal, R., Su, Y.H., Rieke, N., Laina, I., Kalavakonda, N., Bodenstedt, S., et al., 2019 · 1902
Earlier work this paper cites.
Scibert: A pretrained language model for scientific text
Beltagy, I., Lo, K., Cohan, A., 2019 · 1903
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Huang, K., Altosaar, J., Ranganath, R., 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V., 2019 · 1907
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L.H., Yatskar, M., Yin, D., Hsieh, C.J., Chang, K.W., 2019 · 1908
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., Zettlemoyer, L., 2019 · 1910
Earlier work this paper cites.
Temporal memory relation network for workflow recognition from surgical video
Jin, Y., Long, Y., Chen, C., Zhao, Z., Dou, Q., Heng, P.A., 2021 · 1923
Earlier work this paper cites.
Univl: A unified video and language pre-training model for multimodal understanding and generation
Luo, H., Ji, L., Shi, B., Huang, H., Duan, N., Li, T., Li, J., Bharti, T., Zhou, M., 2020 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318
Papineni, K., Roukos, S., Ward, T., Zhu, W.J., 2002 · 2002
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., He, K., 2020b · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, pp. 74–81
Lin, C.Y., 2004 · 2004
Earlier work this paper cites.
Daisi: Database for ai surgical instruction
Rojas-Muñoz, E., Couperus, K., Wachs, J., 2020 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments, in: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72
Banerjee, S., Lavie, A., 2005 · 2005
Earlier work this paper cites.
Modeling and online recognition of surgical phases using hidden markov models, in: Medical Image Computing and Computer-Assisted Intervention–MICCAI 2008: 11th International Conference, New York, NY, USA, September 6-10, 2008, Proceedings, Part II 11, Springer. pp. 627–635
Blum, T., Padoy, N., Feußner, H., Navab, N., 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee. pp. 248–255
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009 · 2009
Earlier work this paper cites.
Modeling and segmentation of surgical workflow from laparoscopic video, in: Medical Image Computing and Computer-Assisted Intervention–MICCAI 2010: 13th International Conference, Beijing, China, September 20-24, 2010, Proceedings, Part III 13, Springer. pp. 400–407
Blum, T., Feußner, H., Navab, N., 2010 · 2010
Earlier work this paper cites.
Feature engineering in context-dependent deep neural networks for conversational speech transcription, in: 2011 IEEE Workshop on Automatic Speech Recognition & Understanding, IEEE. pp. 24–29
Seide, F., Li, G., Chen, X., Yu, D., 2011 · 2011
Earlier work this paper cites.
Statistical modeling and recognition of surgical workflow
Padoy, N., Blum, T., Ahmadi, S.A., Feussner, H., Berger, M.O., Navab, N., 2012 · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A.R., Shah, M., 2012 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., Dean, J., 2013 · 2013
Earlier work this paper cites.
Zero-shot event detection using multi-modal fusion of weakly supervised concepts, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2665–2672
Wu, S., Bondugula, S., Luisier, F., Zhuang, X., Natarajan, P., 2014 · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.Y., Vedantam, R., Gupta, S., Dollár, P., Zitnick, C.L., 2015 · 2015
Earlier work this paper cites.
Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, pp. 1440–1448
Girshick, R., 2015 · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin, in: International conference on machine learning, PMLR. pp. 173–182
Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Cheng, Q., Chen, G., et al., 2016 · 2016
Earlier work this paper cites.
Automatic data-driven real-time segmentation and recognition of surgical workflow
Dergachyova, O., Bouget, D., Huaulmé, A., Morandi, X., Jannin, P., 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
Multimodal video classification with stacked contractive autoencoders
Liu, Y., Feng, X., Zhou, Z., 2016 · 2016
Earlier work this paper cites.
Fusing audio, visual and textual clues for sentiment analysis from multimodal content
Poria, S., Cambria, E., Howard, N., Huang, G.B., Hussain, A., 2016 · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis, in: International conference on machine learning, PMLR. pp. 1060–1069
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., Lee, H., 2016 · 2016
Earlier work this paper cites.
Endonet: a deep architecture for recognition tasks on laparoscopic videos
Twinanda, A.P., Shehata, S., Mutter, D., Marescaux, J., De Mathelin, M., Padoy, N., 2016 · 2016
Earlier work this paper cites.
Anticipating visual representations from unlabeled video, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 98–106
Vondrick, C., Pirsiavash, H., Torralba, A., 2016 · 2016
Earlier work this paper cites.
Tall: Temporal activity localization via language query, in: Proceedings of the IEEE international conference on computer vision, pp. 5267–5275
Gao, J., Sun, C., Yang, Z., Nevatia, R., 2017 · 2017
Earlier work this paper cites.
Exploiting feature and class relationships in video categorization with regularized deep neural networks
Jiang, Y.G., Wu, Z., Wang, J., Xue, X., Chang, S.F., 2017 · 2017
Earlier work this paper cites.
Sv-rcnet: workflow recognition from surgical videos using recurrent convolutional network
Jin, Y., Dou, Q., Chen, H., Yu, L., Qin, J., Fu, C.W., Heng, P.A., 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al., 2017 · 2017
Earlier work this paper cites.
Recurrent topic-transition gan for visual paragraph generation, in: Proceedings of the IEEE international conference on computer vision, pp. 3362–3371
Liang, X., Hu, Z., Zhang, H., Gan, C., Xing, E.P., 2017 · 2017
Cited alongside, same era.
Surgical data science for next-generation interventions
Maier-Hein, L., Vedula, S.S., Speidel, S., Navab, N., Kikinis, R., Park, A., Eisenmann, M., Feussner, H., Forestier, G., Giannarou, S., et al., 2017 · 2017
Cited alongside, same era.
Monitoring tool usage in surgery videos using boosted convolutional and recurrent neural networks
Al Hajj, H., Lamard, M., Conze, P.H., Cochener, B., Quellec, G., 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2018 · 2018
Cited alongside, same era.
Temporal coherence-based self-supervised learning for laparoscopic workflow analysis, in: OR 2.0 Context-Aware Operating Theaters, Computer Assisted Robotic Endoscopy, Springer. pp. 85–93
Learning transferable visual models from natural language supervision, in: International Conference on Machine Learning, PMLR. pp. 8748–8763
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al., 2021 · 2021
Later among the works it cites.
Long-term temporally consistent unpaired video translation from simulated surgical 3d data, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3343–3353
Rivoir, D., Pfeiffer, M., Docea, R., Kolbinger, F., Riediger, C., Weitz, J., Speidel, S., 2021 · 2021
Later among the works it cites.
Computer vision in surgery
Ward, T.M., Mascagni, P., Ban, Y., Rosman, G., Padoy, N., Meireles, O., Hashimoto, D.A., 2021 · 2021
Later among the works it cites.
Class-incremental domain adaptation with smoothing and calibration for surgical report generation, in: Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part IV 24, Springer. pp. 269–278
Xu, M., Islam, M., Lim, C.M., Ren, H., 2021b · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Funke, I., Jenke, A., Mees, S.T., Weitz, J., Speidel, S., Bodenstedt, S., 2018 · 2018
Cited alongside, same era.
Self-supervised spatiotemporal feature learning via video rotation prediction
Jing, L., Yang, X., Liu, J., Tian, Y., 2018 · 2018
Cited alongside, same era.
Exploring the limits of weakly supervised pretraining, in: Proceedings of the European conference on computer vision (ECCV), pp. 181–196
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., Van Der Maaten, L., 2018 · 2018
Cited alongside, same era.
Toward domain-invariant speech recognition via large scale training, in: 2018 IEEE Spoken Language Technology Workshop (SLT), IEEE. pp. 441–447
Narayanan, A., Misra, A., Sim, K.C., Pundak, G., Tripathi, A., Elfeky, M., Haghani, P., Strohman, T., Bacchiani, M., 2018 · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A.v.d., Li, Y., Vinyals, O., 2018 · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features, in: Proceedings of the European conference on computer vision (ECCV), pp. 631–648
Owens, A., Efros, A.A., 2018 · 2018
Cited alongside, same era.
Exploiting the potential of unlabeled endoscopic video data with self-supervised learning
Ross, T., Zimmerer, D., Vemuri, A., Isensee, F., Wiesenfarth, M., Bodenstedt, S., Both, F., Kessler, P., Wagner, M., Müller, B., et al., 2018 · 2018
Cited alongside, same era.
Tracking emerges by colorizing videos, in: Proceedings of the European conference on computer vision (ECCV), pp. 391–408
Vondrick, C., Shrivastava, A., Fathi, A., Guadarrama, S., Murphy, K., 2018 · 2018
Cited alongside, same era.
Surgical workflow anticipation using instrument interaction, in: Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part IV 24, Springer. pp. 615–625
Yuan, K., Holden, M., Gao, S., Lee, W.S., 2021 · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Zellers, R., Lu, X., Hessel, J., Yu, Y., Park, J.S., Cao, J., Farhadi, A., Choi, Y., 2021 · 2021
Later among the works it cites.
Multi-modal masked autoencoders for medical vision-and-language pre-training, in: Medical Image Computing and Computer Assisted Intervention–MICCAI 2022: 25th International Conference, Singapore, September 18–22, 2022, Proceedings, Part V, Springer. pp. 679–689
Chen, Z., Du, Y., Hu, J., Liu, Y., Li, G., Wan, X., Chang, T.H., 2022a · 2022
Later among the works it cites.
When does contrastive visual representation learning work?, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14755–14764
Cole, E., Yang, X., Wilber, K., Mac Aodha, O., Belongie, S., 2022 · 2022
Later among the works it cites.
Biomedical image analysis competitions: The state of current participation practice
Eisenmann, M., Reinke, A., Weru, V., Tizabi, M.D., Isensee, F., Adler, T.J., Godau, P., Cheplygina, V., Kozubek, M., Ali, S., et al., 2022 · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18995–19012
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al., 2022 · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, in: International conference on machine learning, PMLR. pp. 12888–12900
Li, J., Li, D., Xiong, C., Hoi, S., 2022 · 2022
Later among the works it cites.
Surgical data science–from concepts toward clinical translation
Maier-Hein, L., Eisenmann, M., Sarikaya, D., März, K., Collins, T., Malpani, A., Fallert, J., Feussner, H., Giannarou, S., Mascagni, P., et al., 2022 · 2022
Later among the works it cites.
Computer vision in surgery: from potential to clinical value
Mascagni, P., Alapatt, D., Sestini, L., Altieri, M.S., Madani, A., Watanabe, Y., Alseidi, A., Redan, J.A., Alfieri, S., Costamagna, G., et al., 2022 · 2022
Later among the works it cites.
Text-only training for image captioning using noise-injected clip
Nukrai, D., Mokady, R., Globerson, A., 2022 · 2022
Later among the works it cites.
Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos
Nwoye, C.I., Yu, T., Gonzalez, C., Seeliger, B., Mascagni, P., Mutter, D., Marescaux, J., Padoy, N., 2022 · 2022
Later among the works it cites.
Rivoir, D., Funke, I., Speidel, S., 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B., 2022 · 2022
Later among the works it cites.
Clip-forge: Towards zero-shot text-to-shape generation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18603–18613
Sanghi, A., Chu, H., Lambourne, J.G., Wang, Y., Cheng, C.Y., Fumero, M., Malekshan, K.R., 2022 · 2022
Later among the works it cites.
Surgical-vqa: Visual question answering in surgical scenes using transformer, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 33–43
Seenivasan, L., Islam, M., Krishna, A.K., Ren, H., 2022 · 2022
Later among the works it cites.
K-lite: Learning transferable visual models with external knowledge
Shen, S., Li, C., Hu, X., Xie, Y., Yang, J., Zhang, P., Gan, Z., Wang, L., Yuan, L., Liu, C., et al., 2022 · 2022
Later among the works it cites.
Neural rendering for stereo 3d reconstruction of deformable tissues in robotic surgery, in: Medical Image Computing and Computer Assisted Intervention–MICCAI 2022: 25th International Conference, Singapore, September 18–22, 2022, Proceedings, Part VII, Springer. pp. 431–441
Wang, Y., Long, Y., Fan, S.H., Dou, Q., 2022 · 2022
Later among the works it cites.
Rethinking surgical captioning: End-to-end window-based mlp transformer using patches, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 376–386
Xu, M., Islam, M., Ren, H., 2022 · 2022
Later among the works it cites.
Amazon transcribe medical
AWS, 2023 · 2023
Closest in time.
The european association of endoscopic surgery
EAES, 2023 · 2023
Closest in time.
Knowledge-aware prompt tuning for generalizable vision-language models, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15670–15680
Kan, B., Wang, T., Lu, W., Zhen, X., Guan, W., Zheng, F., 2023 · 2023
Closest in time.
A review of deep learning techniques for speech processing
Mehrish, A., Majumder, N., Bhardwaj, R., Poria, S., 2023 · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision, in: International Conference on Machine Learning, PMLR. pp. 28492–28518
Radford, A., Kim, J.W., Xu, T., Brockman, G., McLeavey, C., Sutskever, I., 2023 · 2023
Closest in time.
Dissecting self-supervised learning methods for surgical computer vision
Ramesh, S., Srivastav, V., Alapatt, D., Yu, T., Murali, A., Sestini, L., Nwoye, C.I., Hamoud, I., Sharma, S., Fleurentin, A., et al., 2023 · 2023
Closest in time.
Clip for all things zero-shot sketch-based image retrieval, fine-grained or not, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2765–2775
Sain, A., Bhunia, A.K., Chowdhury, P.N., Koley, S., Xiang, T., Song, Y.Z., 2023 · 2023
Closest in time.
Glsformer: Gated-long, short sequence transformer for step recognition in surgical videos
Shah, N.A., Sikder, S., Vedula, S.S., Patel, V.M., 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al., 2023 · 2023
Closest in time.
Websurg, the online university of ircad
Websurg, 2023 · 2023
Closest in time.
Ra-clip: Retrieval augmented contrastive language-image pre-training, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19265–19274
Xie, C.W., Sun, S., Xiong, X., Zheng, Y., Zhao, D., Zhou, J., 2023 · 2023
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention, in: International conference on machine learning, PMLR. pp. 2048–2057
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y., 2015 · 2057
Closest in time.
Video2vec embeddings recognize events when examples are scarce
Habibian, A., Mensink, T., Snoek, C.G., 2016 · 2089
Closest in time.