Fetching the paper…
Reading the bibliography…
We introduce VidLPRO, a novel video-language (VL) pre-training framework designed specifically for robotic and laparoscopic surgery.
Modeling and online recognition of surgical phases using hidden markov models
Tobias Blum, Nicolas Padoy, Hubertus Feußner, and Nassir Navab · 2008
Earlier work this paper cites.
Modeling and segmentation of surgical workflow from laparoscopic video
Tobias Blum, Hubertus Feußner, and Nassir Navab · 2010
Earlier work this paper cites.
Statistical modeling and recognition of surgical workflow
Nicolas Padoy, Tobias Blum, Seyed-Ahmad Ahmadi, Hubertus Feussner, Marie-Odile Berger, and Nassir Navab · 2012
Earlier work this paper cites.
Jhu-isi gesture and skill assessment working set (jigsaws): A surgical activity dataset for human motion modeling
Yixin Gao, S Swaroop Vedula, Carol E Reiley, Narges Ahmidi, Balakrishnan Varadarajan, Henry C Lin, Lingling Tao, Luca Zappella, Benjamın Béjar, David D Yuh, et al · 2014
Earlier work this paper cites.
Automatic data-driven real-time segmentation and recognition of surgical workflow
Olga Dergachyova, David Bouget, Arnaud Huaulmé, Xavier Morandi, and Pierre Jannin · 2016
Earlier work this paper cites.
Endonet: a deep architecture for recognition tasks on laparoscopic videos
Andru P Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel De Mathelin, and Nicolas Padoy · 2016
Earlier work this paper cites.
Automatic data-driven real-time segmentation and recognition of surgical workflow
Olga Dergachyova, David Bouget, Arnaud Huaulmé, Xavier Morandi, and Pierre Jannin · 2016
Earlier work this paper cites.
Surgical data science for next-generation interventions
Lena Maier-Hein, Swaroop S Vedula, Stefanie Speidel, Nassir Navab, Ron Kikinis, Adrian Park, Matthias Eisenmann, Hubertus Feussner, Germain Forestier, Stamatia Giannarou, et al · 2017
Earlier work this paper cites.
Sv-rcnet: workflow recognition from surgical videos using recurrent convolutional network
Yueming Jin, Qi Dou, Hao Chen, Lequan Yu, Jing Qin, Chi-Wing Fu, and Pheng-Ann Heng · 2017
Earlier work this paper cites.
Comparative evaluation of instrument segmentation and tracking methods in minimally invasive surgery
Sebastian Bodenstedt, Max Allan, Anthony Agustinos, Xiaofei Du, Luis Garcia-Peraza-Herrera, Hannes Kenngott, Thomas Kurmann, Beat Müller-Stich, Sebastien Ourselin, Daniil Pakhomov, et al · 2018
Earlier work this paper cites.
Sv-rcnet: Workflow recognition from surgical videos using recurrent convolutional network
Yueming Jin, Qi Dou, Hao Chen, Lequan Yu, Jing Qin, Chi-Wing Fu, and Pheng-Ann Heng · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Earlier work this paper cites.
2017 robotic instrument segmentation challenge, 2019
Max Allan, Alex Shvets, Thomas Kurmann, Zichen Zhang, Rahul Duggal, Yun-Hsuan Su, Nicola Rieke, Iro Laina, Niveditha Kalavakonda, Sebastian Bodenstedt, Luis Herrera, Wenqi Li, Vladimir Iglovikov, Huoling Luo, Jian Yang, Danail Stoyanov, Lena Maier-Hein, Stefanie Speidel, and Mahdi Azizian · 2019
Earlier work this paper cites.
Generating large labeled data sets for laparoscopic image processing tasks using unpaired image-to-image translation, 2019
Micha Pfeiffer, Isabel Funke, Maria R. Robu, Sebastian Bodenstedt, Leon Strenger, Sandy Engelhardt, Tobias Roß, Matthew J. Clarkson, Kurinchi Gurusamy, Brian R. Davidson, Lena Maier-Hein, Carina Riediger, Thilo Welsch, Jürgen Weitz, and Stefanie Speidel · 2019
Earlier work this paper cites.
Machine and deep learning for workflow recognition during surgery
Nicolas Padoy · 2019
Earlier work this paper cites.
Cai4cai: the rise of contextual artificial intelligence in computer-assisted interventions
Tom Vercauteren, Mathias Unberath, Nicolas Padoy, and Nassir Navab · 2019
Earlier work this paper cites.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language, 2019
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou · 2019
Earlier work this paper cites.
VL-BERT: pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Earlier work this paper cites.
Multi-task recurrent convolutional network with correlation loss for surgical video analysis, 2019
Yueming Jin, Huaxia Li, Qi Dou, Hao Chen, Jing Qin, Chi-Wing Fu, and Pheng-Ann Heng · 2019
Earlier work this paper cites.
Uniter: Universal image-text representation learning, 2020
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti · 2020
Cited alongside, same era.
Tecno: Surgical phase recognition with multi-stage temporal convolutional networks
Tobias Czempiel, Magdalini Paschali, Matthias Keicher, Walter Simson, Hubertus Feussner, Seong Tae Kim, and Nassir Navab · 2020
Cited alongside, same era.
Multi-task recurrent convolutional network with correlation loss for surgical video analysis
Yueming Jin, Huaxia Li, Qi Dou, Hao Chen, Jing Qin, Chi-Wing Fu, and Pheng-Ann Heng · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Neural rendering for stereo 3d reconstruction of deformable tissues in robotic surgery, 2022
Yuehao Wang, Yonghao Long, Siu Hin Fan, and Qi Dou · 2022
Later among the works it cites.
Autolaparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hysterectomy
Ziyi Wang, Bo Lu, Yonghao Long, Fangxun Zhong, Tak-Hong Cheung, Qi Dou, and Yunhui Liu · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Later among the works it cites.
Surgical-vqa: Visual question answering in surgical scenes using transformer, 2022
Lalithkumar Seenivasan, Mobarakol Islam, Adithya K Krishna, and Hongliang Ren · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision, 2022
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Cited alongside, same era.
Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos
Chinedu Innocent Nwoye, Tong Yu, Cristians Gonzalez, Barbara Seeliger, Pietro Mascagni, Didier Mutter, Jacques Marescaux, and Nicolas Padoy · 2021
Cited alongside, same era.
Temporally constrained neural networks (tcnn): A framework for semi-supervised video semantic segmentation, 2021
Deepak Alapatt, Pietro Mascagni, Armine Vardazaryan, Alain Garcia, Nariaki Okamoto, Didier Mutter, Jacques Marescaux, Guido Costamagna, Bernard Dallemagne, and Nicolas Padoy · 2021
Cited alongside, same era.
Long-term temporally consistent unpaired video translation from simulated surgical 3d data, 2021
Dominik Rivoir, Micha Pfeiffer, Reuben Docea, Fiona Kolbinger, Carina Riediger, Jürgen Weitz, and Stefanie Speidel · 2021
Cited alongside, same era.
Surgical visual domain adaptation: Results from the miccai 2020 surgvisdom challenge
Aneeq Zia, Kiran Bhattacharyya, Xi Liu, Ziheng Wang, Satoshi Kondo, Emanuele Colleoni, Beatrice van Amsterdam, Razeen Hussain, Raabid Hussain, Lena Maier-Hein, et al · 2021
Cited alongside, same era.
VIOLET: End-to-End Video-Language Transformers with Masked Visual-token Modeling
Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu · 2021
Cited alongside, same era.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Cited alongside, same era.
Omnivl:one foundation model for image-language and video-language tasks, 2022
Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan · 2022
Later among the works it cites.
Revealing single frame bias for video-and-language learning, 2022
Jie Lei, Tamara L. Berg, and Mohit Bansal · 2022
Later among the works it cites.
Multimodal dataset distillation for image-text retrieval
Xindi Wu, Byron Zhang, Zhiwei Deng, and Olga Russakovsky · 2023
Later among the works it cites.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al · 2023
Later among the works it cites.
Lavender: Unifying video-language understanding as masked language modeling
Linjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin, Ce Liu, Zicheng Liu, and Lijuan Wang · 2023
Later among the works it cites.
Learning multi-modal representations by watching hundreds of surgical video lectures
Kun Yuan, Vinkle Srivastav, Tong Yu, Joel Lavanchy, Pietro Mascagni, Nassir Navab, and Nicolas Padoy · 2023
Later among the works it cites.
Surgicalgpt: End-to-end language-vision gpt for visual question answering in surgery, 2023
Lalithkumar Seenivasan, Mobarakol Islam, Gokul Kannan, and Hongliang Ren · 2023
Later among the works it cites.
Lovit: Long video transformer for surgical phase recognition, 2023
Yang Liu, Maxence Boels, Luis C. Garcia-Peraza-Herrera, Tom Vercauteren, Prokar Dasgupta, Alejandro Granados, and Sebastien Ourselin · 2023
Later among the works it cites.
Skit: a fast key information video transformer for online surgical phase recognition
Yang Liu, Jiayu Huo, Jingjing Peng, Rachel Sparks, Prokar Dasgupta, Alejandro Granados, and Sebastien Ourselin · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Vindlu: A recipe for effective video-and-language pretraining
Feng Cheng, Xizi Wang, Jie Lei, David Crandall, Mohit Bansal, and Gedas Bertasius · 2023
Later among the works it cites.
Videoprism: A foundational visual encoder for video understanding, 2024
Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J. Sun, Luke Friedman, Rui Qian, Tobias Weyand, Yue Zhao, Rachel Hornung, Florian Schroff, Ming-Hsuan Yang, David A. Ross, Huisheng Wang, Hartwig Adam, Mikhail Sirotenko, Ting Liu, and Boqing Gong · 2024
Closest in time.
Hecvl: Hierarchical video-language pretraining for zero-shot surgical phase recognition
Kun Yuan, Vinkle Srivastav, Nassir Navab, and Nicolas Padoy · 2024
Closest in time.
General surgery vision transformer: A video pre-trained foundation model for general surgery
Samuel Schmidgall, Ji Woong Kim, Jeffery Jopling, and Axel Krieger · 2024
Closest in time.
Surgical-lvlm: Learning to adapt large vision-language model for grounded visual question answering in robotic surgery, 2024
Guankun Wang, Long Bai, Wan Jun Nah, Jie Wang, Zhaoxi Zhang, Zhen Chen, Jinlin Wu, Mobarakol Islam, Hongbin Liu, and Hongliang Ren · 2024
Closest in time.