Fetching the paper…
Reading the bibliography…
Surgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data.
Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer
Juan Silva, Aymeric Histace, Olivier Romain, Xavier Dray, and Bertrand Granado · 2014
Earlier work this paper cites.
Wm-dova maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians
Jorge Bernal, F Javier Sánchez, Gloria Fernández-Esparrach, Debora Gil, Cristina Rodríguez, and Fernando Vilariño · 2015
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Concurrent segmentation and localization for tracking of surgical instruments
Iro Laina, Nicola Rieke, Christian Rupprecht, Josué Page Vizcaíno, Abouzar Eslami, Federico Tombari, and Nassir Navab · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Frame-based classification of operation phases in cataract surgery videos
Manfred Jürgen Primus, Doris Putzgruber-Adamitsch, Mario Taschwer, Bernd Münzer, Yosuf El-Shabrawi, László Böszörményi, and Klaus Schoeffmann · 2018
Earlier work this paper cites.
Automatic instrument segmentation in robot-assisted surgery using deep learning
Alexey A Shvets, Alexander Rakhlin, Alexandr A Kalinin, and Vladimir I Iglovikov · 2018
Earlier work this paper cites.
Reconstruction network for video captioning
Bairui Wang, Lin Ma, Wei Zhang, and Wei Liu · 2018
Earlier work this paper cites.
Cross-modal and hierarchical modeling of video and text
Bowen Zhang, Hexiang Hu, and Fei Sha · 2018
Earlier work this paper cites.
End-to-end dense video captioning with masked transformer
Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong · 2018
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath · 2019
Earlier work this paper cites.
Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs
Alistair EW Johnson, Tom J Pollard, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Yifan Peng, Zhiyong Lu, Roger G Mark, Seth J Berkowitz, and Steven Horng · 2019
Earlier work this paper cites.
Deep residual learning for instrument segmentation in robotic surgery
Daniil Pakhomov, Vittal Premachandran, Max Allan, Mahdi Azizian, and Nassir Navab · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
Pixel-based tool segmentation in cataract surgery videos with mask R-CNN
Markus Fox, Mario Taschwer, and Klaus Schoeffmann · 2020
Earlier work this paper cites.
Relevance detection in cataract surgery videos by spatio- temporal action localization
Negin Ghamsarian, Mario Taschwer, Doris Putzgruber-Adamitsch, Stephanie Sarny, and Klaus Schoeffmann · 2020
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing, 2020
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2020
Earlier work this paper cites.
Towards unsupervised learning for instrument segmentation in robotic surgery with cycle-consistent adversarial networks
Daniil Pakhomov, Wei Shen, and Nassir Navab · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Earlier work this paper cites.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2021
Earlier work this paper cites.
Lensid: A cnn-rnn-based framework towards lens irregularity detection in cataract surgery videos
Negin Ghamsarian, Mario Taschwer, Doris Putzgruber-Adamitsch, Stephanie Sarny, Yosuf El-Shabrawi, and Klaus Schoeffmann · 2021
Earlier work this paper cites.
Cadis: Cataract dataset for surgical rgb-image segmentation
Maria Grammatikopoulou, Evangello Flouty, Abdolrahim Kadkhodamohammadi, Gwenolé Quellec, Andre Chow, Jean Nehme, Imanol Luengo, and Danail Stoyanov · 2021
Earlier work this paper cites.
A comprehensive study on colorectal polyp segmentation with resunet++, conditional random field and test-time augmentation
Debesh Jha, Pia H Smedsrud, Dag Johansen, Thomas De Lange, Håvard D Johansen, Pål Halvorsen, and Michael A Riegler · 2021
Earlier work this paper cites.
Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Kvasir-Capsule, a video capsule endoscopy dataset
Pia H Smedsrud, Vajira Thambawita, Steven A Hicks, Henrik Gjestang, Oda Olsen Nedrejord, Espen Næss, Hanna Borgli, Debesh Jha, Tor Jan Derek Berstad, Sigrun L Eskeland, Mathias Lux, Håvard Espeland, Andreas Petlund, Duc Tien Dang Nguyen, Enrique Garcia-Ceja, Dag Johansen, Peter T Schmidt, Ervin Toth, Hugo L Hammer, Thomas de Lange, Michael A Riegler, and Pål Halvorsen · 2021
Cited alongside, same era.
Multimodal contrastive training for visual representation learning
Xin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang, Yilin Wang, Michael Maire, Ajinkya Kale, and Baldo Faieta · 2021
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2022
Cited alongside, same era.
Real-time instance segmentation of surgical instruments using attention and multi-scale feature fusion
Capabilities of gpt-4 on medical challenge problems
Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz · 2023
Later among the works it cites.
Artificial intelligence in surgical learning
Niklas Pakkasjärvi, Tanvi Luthra, and Sachit Anand · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
In-context retrieval-augmented language models
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham · 2023
Later among the works it cites.
Smallcap: lightweight image captioning prompted with retrieval augmentation
Rita Ramos, Bruno Martins, Desmond Elliott, and Yova Kementchedjhieva · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Juan Carlos Ángeles Cerón, Gilberto Ochoa Ruiz, Leonardo Chang, and Sharib Ali · 2022
Cited alongside, same era.
Retrieval augmented classification for long-tail visual recognition
Alexander Long, Wei Yin, Thalaiyasingam Ajanthan, Vu Nguyen, Pulak Purkait, Ravi Garg, Alan Blair, Chunhua Shen, and Anton van den Hengel · 2022
Cited alongside, same era.
Image segmentation using text and image prompts
Timo Lüddecke and Alexander Ecker · 2022
Cited alongside, same era.
Upolyseg: A u-net-based polyp segmentation network using colonoscopy images
Subhashree Mohapatra, Girish Kumar Pati, Manohar Mishra, and Tripti Swarnkar · 2022
Cited alongside, same era.
Slip: Self-supervision meets language-image pre-training
Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie · 2022
Cited alongside, same era.
Expanding language-image pretrained models for general video recognition
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling · 2022
Cited alongside, same era.
Medical image understanding with pretrained vision language models: A comprehensive study
Ziyuan Qin, Huahui Yi, Qicheng Lao, and Kang Li · 2022
Cited alongside, same era.
Retrieval-augmented transformer for image captioning
Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara · 2022
Cited alongside, same era.
Madeline C Schiappa, Yogesh S Rawat, and Mubarak Shah · 2023
Later among the works it cites.
Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery
Lalithkumar Seenivasan, Mobarakol Islam, Gokul Kannan, and Hongliang Ren · 2023
Later among the works it cites.
Comparative validation of machine learning algorithms for surgical workflow and skill analysis with the heichole benchmark
Martin Wagner, Beat-Peter Müller-Stich, Anna Kisilenko, Duc Tran, Patrick Heger, Lars Mündermann, David M Lubotsky, Benjamin Müller, Tornike Davitashvili, Manuela Capek, et al · 2023
Later among the works it cites.
Learning multi-modal representations by watching hundreds of surgical video lectures
Kun Yuan, Vinkle Srivastav, Tong Yu, Joel L Lavanchy, Pietro Mascagni, Nassir Navab, and Nicolas Padoy · 2023
Later among the works it cites.
Huatuogpt, towards taming language model to be a doctor
Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, et al · 2023
Later among the works it cites.
Text promptable surgical instrument segmentation with vision-language models
Zijian Zhou, Oluwatosin Alabi, Meng Wei, Tom Vercauteren, and Miaojing Shi · 2023
Later among the works it cites.
Generalized decoding for pixel, image, and language
Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, et al · 2023
Later among the works it cites.
Chexagent: Towards a foundation model for chest x-ray interpretation
Zhihong Chen, Maya Varma, Jean-Benoit Delbrouck, Magdalini Paschali, Louis Blankemeier, Dave Van Veen, Jeya Maria Jose Valanarasu, Alaa Youssef, Joseph Paul Cohen, Eduardo Pontes Reis, et al · 2024
Closest in time.
Improving clip training with language rewrites
Lijie Fan, Dilip Krishnan, Phillip Isola, Dina Katabi, and Yonglong Tian · 2024
Closest in time.
Cataract-1k dataset for deep-learning-assisted analysis of cataract surgery videos
Negin Ghamsarian, Yosuf El-Shabrawi, Sahar Nasirihaghighi, Doris Putzgruber-Adamitsch, Martin Zinkernagel, Sebastian Wolf, Klaus Schoeffmann, and Raphael Sznitman · 2024
Closest in time.
Vidlpro: A video-language pre-training framework for robotic and laparoscopic surgery
Mohammadmahdi Honarmand, Muhammad Abdullah Jamal, and Omid Mohareri · 2024
Closest in time.
Ophnet: A large-scale video benchmark for ophthalmic surgical workflow understanding
Ming Hu, Peng Xia, Lin Wang, Siyuan Yan, Feilong Tang, Zhongxing Xu, Yimin Luo, Kaimin Song, Jurgen Leitner, Xuelian Cheng, et al · 2024
Closest in time.
Quilt-1m: One million image-text pairs for histopathology
Wisdom Ikezogwo, Saygin Seyfioglu, Fatemeh Ghezloo, Dylan Geva, Fatwir Sheikh Mohammed, Pavan Kumar Anand, Ranjay Krishna, and Linda Shapiro · 2024
Closest in time.
Self-supervised learning of semantic correspondence using web videos
Donghyeon Kwon, Minsu Cho, and Suha Kwak · 2024
Closest in time.
Artificial intelligence in surgery
Chris Varghese, Ewen M Harrison, Greg O’Grady, and Eric J Topol · 2024
Closest in time.
Guankun Wang, Long Bai, Wan Jun Nah, Jie Wang, Zhaoxi Zhang, Zhen Chen, Jinlin Wu, Mobarakol Islam, Hongbin Liu, and Hongliang Ren · 2024
Closest in time.
Retrieval-augmented egocentric video captioning
Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie · 2024
Closest in time.
Segment everything everywhere all at once
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee · 2024
Closest in time.