Fetching the paper…
Reading the bibliography…
We find Mask2Former also achieves state-of-the-art performance on video instance segmentation without modifying the architecture, the loss or even the training pipeline.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Mots: Multi-object tracking and segmentation
Paul Voigtlaender, Michael Krause, Aljosa Osep, Jonathon Luiten, Berin Balachandar Gnana Sekar, Andreas Geiger, and Bastian Leibe · 2019
Earlier work this paper cites.
Detectron2
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick · 2019
Earlier work this paper cites.
Video instance segmentation
Linjie Yang, Yuchen Fan, and Ning Xu · 2019
Cited alongside, same era.
Stem-seg: Spatio-temporal embeddings for instance segmentation in videos
Ali Athar, Sabarinath Mahadevan, Aljosa Osep, Laura Leal-Taixé, and Bastian Leibe · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Panoptic-DeepLab: A simple, strong, and fast baseline for bottom-up panoptic segmentation
Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen · 2020
Cited alongside, same era.
Video panoptic segmentation
Dahun Kim, Sanghyun Woo, Joon-Young Lee, and In So Kweon · 2020
Cited alongside, same era.
Masked-attention mask transformer for universal image segmentation
Per-pixel classification is not all you need for semantic segmentation
Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov · 2021
Closest in time.
Video instance segmentation using inter-frame communication transformers
Sukjun Hwang, Miran Heo, Seoung Wug Oh, and Seon Joo Kim · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Closest in time.
Vip-deeplab: Learning visual perception with depth-aware video panoptic segmentation
Siyuan Qiao, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen · 2021
Closest in time.
End-to-end video instance segmentation with transformers
Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar · 2021
Cited alongside, same era.
Junfeng Wu, Yi Jiang, Wenqing Zhang, Xiang Bai, and Song Bai · 2021
Closest in time.