Fetching the paper…
Reading the bibliography…
Until recently, the Video Instance Segmentation (VIS) community operated under the common belief that offline methods are generally superior to a frame by frame online processing.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., and Dollár, P · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Youtube-vos: A large-scale video object segmentation benchmark
Xu, N., Yang, L., Fan, Y., Yue, D., Liang, Y., Yang, J., and Huang, T. S · 2018
Earlier work this paper cites.
Video instance segmentation
Yang, L., Fan, Y., and Xu, N · 2019
Earlier work this paper cites.
Stem-seg: Spatio-temporal embeddings for instance segmentation in videos
Athar, A., Mahadevan, S., Ošep, A., Leal-Taixé, L., and Leibe, B · 2020
Earlier work this paper cites.
Classifying, segmenting, and tracking object instances in video with mask propagation
Bertasius, G. and Torresani, L · 2020
Earlier work this paper cites.
Sipmask: Spatial information preservation for fast image and video instance segmentation
Cao, J., Anwer, R. M., Cholakkal, H., Khan, F. S., Pang, Y., and Shao, L · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Earlier work this paper cites.
PointRend: Image segmentation as rendering
Kirillov, A., Wu, Y., He, K., and Girshick, R · 2020
Cited alongside, same era.
Mask2former for video instance segmentation
Cheng, B., Choudhuri, A., Misra, I., Kirillov, A., Girdhar, R., and Schwing, A. G · 2021
Cited alongside, same era.
Compfeat: Comprehensive feature aggregation for video instance segmentation
Fu, Y., Yang, L., Liu, D., Huang, T. S., and Shi, H · 2021
Cited alongside, same era.
Video instance segmentation using inter-frame communication transformers
Hwang, S., Heo, M., Oh, S. W., and Kim, S. J · 2021
Cited alongside, same era.
Spatial feature calibration and temporal fusion for effective one-stage video instance segmentation
Li, M., Li, S., Li, L., and Zhang, L · 2021
Cited alongside, same era.
Video instance segmentation with a propose-reduce paradigm
Masked-attention mask transformer for universal image segmentation
Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., and Girdhar, R · 2022
Later among the works it cites.
VISOLO: grid-based space-time aggregation for efficient online video instance segmentation
Han, S. H., Hwang, S., Oh, S. W., andHyunwoo Kim, Y. P., Kim, M., and Kim, S. J · 2022
Later among the works it cites.
Vita: Video instance segmentation via object token association
Heo, M., Hwang, S., Oh, S. W., Lee, J.-Y., and Kim, S. J · 2022
Later among the works it cites.
Minvis: A minimal video instance segmentation framework without video-based training
Huang, D.-A., Yu, Z., and Anandkumar, A · 2022
Later among the works it cites.
Occluded video instance segmentation: A benchmark
Qi, J., Gao, Y., Hu, Y., Wang, X., Liu, X., Bai, X., Belongie, S., Yuille, A., Torr, P., and Bai, S · 2022
Later among the works it cites.
Temporally efficient vision transformer for video instance segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, H., Wu, R., Liu, S., Lu, J., and Jia, J · 2021
Cited alongside, same era.
End-to-end video instance segmentation with transformers
Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., and Xia, H · 2021
Cited alongside, same era.
Crossover learning for fast online video instance segmentation
Yang, S., Fang, Y., Wang, X., Li, Y., Fang, C., Shan, Y., Feng, B., and Liu, W · 2021
Cited alongside, same era.
DeVIS: Making Deformable Transformers Work for Video Instance Segmentation
Caelles, A., Meinhardt, T., Brasó, G., and Leal-Taixé, L · 2022
Cited alongside, same era.
Sg-net: Spatial granularity network for one-stage video instance segmentation
Liu, D., Cui, Y., Tan, W., and Chen, Y
Cited in the paper.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B
Cited in the paper.
Seqformer: Sequential transformer for video instance segmentation
Wu, J., Jiang, Y., Bai, S., Zhang, W., and Bai, X
Cited in the paper.
Yang, S., Wang, X., Li, Y., Fang, Y., Fang, J., Liu, Zhao, X., and Shan, Y · 2022
Later among the works it cites.
TransVOD: End-to-end Video Object Detection with Spatial-Temporal Transformers
Zhou, Q., Li, X., He, L., Yang, Y., Cheng, G., Tong, Y., Ma, L., and Tao, D · 2022
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J · 2022
Later among the works it cites.