Fetching the paper…
Reading the bibliography…
We introduce a novel paradigm for offline Video Instance Segmentation (VIS), based on the hypothesis that explicit object-oriented information can be a strong clue for understanding the context of the entire sequence.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
Global data association for multi-object tracking using network flows
Li Zhang, Yuan Li, and Ramakant Nevatia · 2008
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick · 2017
Earlier work this paper cites.
Multiple people tracking by lifted multicut and person re-identification
Siyu Tang, Mykhaylo Andriluka, Bjoern Andres, and Bernt Schiele · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Detectron2
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick · 2019
Earlier work this paper cites.
Video instance segmentation
Linjie Yang, Yuchen Fan, and Ning Xu · 2019
Earlier work this paper cites.
Stem-seg: Spatio-temporal embeddings for instance segmentation in videos
Ali Athar, Sabarinath Mahadevan, Aljoša Ošep, Laura Leal-Taixé, and Bastianan Leibe · 2020
Earlier work this paper cites.
Classifying, segmenting, and tracking object instances in video with mask propagation
Gedas Bertasius and Lorenzo Torresani · 2020
Earlier work this paper cites.
Learning a neural solver for multiple object tracking
Guillem Brasó and Laura Leal-Taixé · 2020
Earlier work this paper cites.
Sipmask: Spatial information preservation for fast image and video instance segmentation
Jiale Cao, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, and Ling Shao · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
End-to-end video instance segmentation with transformers
Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia · 2020
Cited alongside, same era.
Mask2former for video instance segmentation
Bowen Cheng, Anwesa Choudhuri, Ishan Misra, Alexander Kirillov, Rohit Girdhar, and Alexander G Schwing · 2021
Cited alongside, same era.
Learning a proposal classifier for multiple object tracking
Peng Dai, Renliang Weng, Wongun Choi, Changshui Zhang, Zhangping He, and Wei Ding · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Cited alongside, same era.
Occluded video instance segmentation
Jiyang Qi, Yan Gao, Yao Hu, Xinggang Wang, Xiaoyu Liu, Xiang Bai, Serge Belongie, Alan Yuille, Philip HS Torr, and Song Bai · 2021
Later among the works it cites.
Occluded video instance segmentation: Dataset and iccv 2021 challenge
Jiyang Qi, Yan Gao, Yao Hu, Xinggang Wang, Xiaoyu Liu, Xiang Bai, Serge Belongie, Alan Yuille, Philip HS Torr, and Song Bai · 2021
Later among the works it cites.
Rethinking transformer-based set prediction for object detection
Zhiqing Sun, Shengcao Cao, Yiming Yang, and Kris M Kitani · 2021
Later among the works it cites.
Crossover learning for fast online video instance segmentation
Shusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu · 2021
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Compfeat: Comprehensive feature aggregation for video instance segmentation
Yang Fu, Linjie Yang, Ding Liu, Thomas S Huang, and Humphrey Shi · 2021
Cited alongside, same era.
Fast convergence of detr with spatially modulated co-attention
Peng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai, and Hongsheng Li · 2021
Cited alongside, same era.
Video instance segmentation using inter-frame communication transformers
Sukjun Hwang, Miran Heo, Seoung Wug Oh, and Seon Joo Kim · 2021
Cited alongside, same era.
Prototypical cross-attention networks for multiple object tracking and segmentation
Lei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, and Fisher Yu · 2021
Cited alongside, same era.
Spatial feature calibration and temporal fusion for effective one-stage video instance segmentation
Minghan Li, Shuai Li, Lida Li, and Lei Zhang · 2021
Cited alongside, same era.
Sg-net: Spatial granularity network for one-stage video instance segmentation
Dongfang Liu, Yiming Cui, Wenbo Tan, and Yingjie Chen · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Closest in time.
Visolo: Grid-based space-time aggregation for efficient online video instance segmentation
Su Ho Han, Sukjun Hwang, Seoung Wug Oh, Yeonchool Park, Hyunwoo Kim, Min-Jung Kim, and Seon Joo Kim · 2022
Closest in time.
Cannot see the forest for the trees: Aggregating multiple viewpoints to better classify objects in videos
Sukjun Hwang, Miran Heo, Seoung Wug Oh, and Seon Joo Kim · 2022
Closest in time.
Efficient video instance segmentation via tracklet query and proposal
Jialian Wu, Sudhir Yarram, Hui Liang, Tian Lan, Junsong Yuan, Jayan Eledath, and Gerard Medioni · 2022
Closest in time.
Seqformer: a frustratingly simple model for video instance segmentation
Junfeng Wu, Yi Jiang, Wenqing Zhang, Xiang Bai, and Song Bai · 2022
Closest in time.
Temporally efficient vision transformer for video instance segmentation
Shusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang, Jiemin Fang, Wenyu Liu, Xun Zhao, and Ying Shan · 2022
Closest in time.
Global tracking transformers
Xingyi Zhou, Tianwei Yin, Vladlen Koltun, and Phillip Krähenbühl · 2022
Closest in time.