Fetching the paper…
Reading the bibliography…
Endoscopic video-based tasks, such as visual navigation and surgical phase recognition, play a crucial role in minimally invasive surgeries by providing real-time assistance.
Bernal, J., Sánchez, F.J., Fernández-Esparrach, G., Gil, D., Rodríguez, C., Vilariño, F.: WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians. Computerized medical imaging and graphics 43
2015
Earlier work this paper cites.
Mesejo, P., Pizarro, D., Abergel, A., Rouquette, O., Beorchia, S., Poincloux, L., Bartoli, A.: Computer-aided classification of gastrointestinal lesions in regular colonoscopy. IEEE transactions on medical imaging 35
2016
Earlier work this paper cites.
2019
Earlier work this paper cites.
Borgli, H., Thambawita, V., Smedsrud, P.H., Hicks, S., Jha, D., Eskeland, S.L., Randel, K.R., Pogorelov, K., Lux, M., Nguyen, D.T.D., et al.: HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy. Scientific data 7
2020
Earlier work this paper cites.
Leibetseder, A., Kletz, S., Schoeffmann, K., Keckstein, S., Keckstein, J.: GLENDA: gynecologic laparoscopy endometriosis dataset. In: MMM. Lecture Notes in Computer Science, vol. 11962, pp. 439–450. Springer (2020)
2020
Earlier work this paper cites.
Gao, X., Jin, Y., Long, Y., Dou, Q., Heng, P.A.: Trans-SVNet: Accurate phase recognition from surgical videos via hybrid embedding aggregation transformer. In: MICCAI. pp. 593–603. Springer (2021)
2021
Earlier work this paper cites.
Girdhar, R., Grauman, K.: Anticipative video transformer. In: ICCV. pp. 13505–13515 (2021)
2021
Earlier work this paper cites.
Ma, Y., Chen, X., Cheng, K., Li, Y., Sun, B.: LDPolypVideo benchmark: a large-scale colonoscopy video dataset of diverse polyps. In: MICCAI. pp. 387–396. Springer (2021)
2021
Earlier work this paper cites.
Misawa, M., Kudo, S.e., Mori, Y., Hotta, K., Ohtsuka, K., Matsuda, T., Saito, S., Kudo, T., Baba, T., Ishida, F., et al.: Development of a computer-aided detection system for colonoscopy and a publicly accessible large colonoscopy video database (with video). Gastrointestinal endoscopy 93
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: ICML. pp. 8748–8763. PMLR (2021)
2021
Earlier work this paper cites.
Roß, T., Reinke, A., Full, P.M., Wagner, M., Kenngott, H., Apitz, M., Hempe, H., Mindroc-Filimon, D., Scholz, P., Tran, T.N., et al.: Comparative validation of multi-instance instrument segmentation in endoscopy: results of the robust-mis 2019 challenge. Medical image analysis 70
2021
Earlier work this paper cites.
Smedsrud, P.H., Thambawita, V., Hicks, S.A., Gjestang, H., Nedrejord, O.O., Næss, E., Borgli, H., Jha, D., Berstad, T.J.D., Eskeland, S.L., et al.: Kvasir-Capsule, a video capsule endoscopy dataset. Scientific Data 8
2021
Earlier work this paper cites.
Yoon, J., Lee, J., Heo, S., Yu, H., Lim, J., Song, C.H., Hong, S., Hong, S., Park, B., Park, S., et al.: hsdb-instrument: instrument localization database for laparoscopic and robotic surgeries. In: MICCAI. pp. 393–402. Springer (2021)
2021
Cited alongside, same era.
Ding, S., Li, M., Yang, T., Qian, R., Xu, H., Chen, Q., Wang, J., Xiong, H.: Motion-aware contrastive video representation learning via foreground-background merging. In: ICCV. pp. 9716–9726 (2022)
2022
Cited alongside, same era.
Ji, G.P., Xiao, G., Chou, Y.C., Fan, D.P., Zhao, K., Chen, G., Van Gool, L.: Video polyp segmentation: A deep learning perspective. Machine Intelligence Research 19
2022
Cited alongside, same era.
Nwoye, C.I., Yu, T., Gonzalez, C., Seeliger, B., Mascagni, P., Mutter, D., Marescaux, J., Padoy, N.: Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos. Medical Image Analysis 78
2022
Cited alongside, same era.
Liu, Y., Huo, J., Peng, J., Sparks, R., Dasgupta, P., Granados, A., Ourselin, S.: SKiT: a fast key information video transformer for online surgical phase recognition. In: ICCV. pp. 21074–21084 (2023)
2023
Later among the works it cites.
Wang, L., Huang, B., Zhao, Z., Tong, Z., He, Y., Wang, Y., Wang, Y., Qiao, Y.: VideoMAE v2: Scaling video masked autoencoders with dual masking. In: ICCV. pp. 14549–14560 (2023)
2023
Later among the works it cites.
Wang, Z., Liu, C., Zhang, S., Dou, Q.: Foundation model for endoscopy video analysis via large-scale self-supervised pre-train. In: MICCAI. pp. 101–111 (2023)
2023
Later among the works it cites.
Dao, T., Gu, A.: Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In: ICML (2024)
2024
Later among the works it cites.
Gu, A., Dao, T.: Mamba: Linear-time sequence modeling with selective state spaces. In: First Conference on Language Modeling (2024)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pan, J., Lin, Z., Zhu, X., Shao, J., Li, H.: ST-Adapter: Parameter-efficient image-to-video transfer learning. NeurIPS 35
2022
Cited alongside, same era.
Park, J., Lee, J., Kim, I.J., Sohn, K.: Probabilistic representations for video contrastive learning. In: ICCV. pp. 14711–14721 (2022)
2022
Cited alongside, same era.
Qian, R., Ding, S., Liu, X., Lin, D.: Static and dynamic concepts for self-supervised video representation learning. In: ECCV. pp. 145–164. Springer (2022)
2022
Cited alongside, same era.
Tian, Y., Pang, G., Liu, F., Liu, Y., Wang, C., Chen, Y., Verjans, J., Carneiro, G.: Contrastive transformer-based multiple instance learning for weakly supervised polyp frame detection. In: MICCAI. pp. 88–98. Springer (2022)
2022
Cited alongside, same era.
Tong, Z., Song, Y., Wang, J., Wang, L.: VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. NeurIPS 35
2022
Cited alongside, same era.
Wang, Z., Lu, B., Long, Y., Zhong, F., Cheung, T.H., Dou, Q., Liu, Y.: Autolaparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hysterectomy. In: MICCAI. pp. 486–496. Springer (2022)
2022
Cited alongside, same era.
Azagra, P., Sostres, C., Ferrández, Á., Riazuelo, L., Tomasini, C., Barbed, O.L., Morlana, J., Recasens, D., Batlle, V.M., Gómez-Rodríguez, J.J., et al.: Endomapper dataset of complete calibrated endoscopy procedures. Scientific Data 10
2023
Cited alongside, same era.
Li, K., Wang, Y., Li, Y., Wang, Y., He, Y., Wang, L., Qiao, Y.: Unmasked teacher: Towards training-efficient video foundation models. In: ICCV. pp. 19948–19960 (2023)
2023
Cited alongside, same era.
2024
Later among the works it cites.
Li, K., Li, X., Wang, Y., He, Y., Wang, Y., Wang, L., Qiao, Y.: Videomamba: State space model for efficient video understanding. In: ECCV. pp. 237–255. Springer (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
Tian, Q., Liao, H., Huang, X., Yang, B., Wu, J., Chen, J., Li, L., Liu, H.: BronchoTrack: Airway lumen tracking for branch-level bronchoscopic localization. IEEE Transactions on Medical Imaging (2024)
2024
Later among the works it cites.
Wang, Y., Li, K., Li, X., Yu, J., He, Y., Chen, G., Pei, B., Zheng, R., Wang, Z., Shi, Y., et al.: InternVideo2: Scaling foundation models for multimodal video understanding. In: ECCV. pp. 396–416. Springer (2024)
2024
Later among the works it cites.
Zhang, K., Zhou, R., Adhikarla, E., Yan, Z., Liu, Y., Yu, J., Liu, Z., Chen, X., Davison, B.D., Ren, H., et al.: A generalist vision–language foundation model for diverse biomedical tasks. Nature Medicine pp. 1–13 (2024)
2024
Later among the works it cites.
Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., Wang, X.: Vision Mamba: Efficient visual representation learning with bidirectional state space model. In: ICML (2024)
2024
Later among the works it cites.
Liu, Y., Boels, M., Garcia-Peraza-Herrera, L.C., Vercauteren, T., Dasgupta, P., Granados, A., Ourselin, S.: LoViT: Long video transformer for surgical phase recognition. Medical Image Analysis 99
2025
Closest in time.