Fetching the paper…
Reading the bibliography…
Video Instance Segmentation (VIS) is a task that simultaneously requires classification, segmentation, and instance association in a video.
Van der Maaten, L., Hinton, G.: Visualizing data using t-sne. Journal of Machine Learning Research 9
2008
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European Conference on Computer Vision. pp. 740–755 (2014)
2014
Earlier work this paper cites.
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al.: Imagenet large scale visual recognition challenge. International Journal of Computer Vision 115
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 770–778 (2016)
2016
Earlier work this paper cites.
Jia, X., De Brabandere, B., Tuytelaars, T., Gool, L.V.: Dynamic filter networks. Advances in Neural Information Processing Systems pp. 667–675 (2016)
2016
Earlier work this paper cites.
Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., Wei, Y.: Deformable convolutional networks. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 764–773 (2017)
2017
Earlier work this paper cites.
He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask r-cnn. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 2961–2969 (2017)
2017
Earlier work this paper cites.
Lin, T.Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2117–2125 (2017)
2017
Earlier work this paper cites.
Zhu, L., Xu, Z., Yang, Y.: Bidirectional multirate reconstruction for temporal modeling in videos. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2653–2662 (2017)
2017
Earlier work this paper cites.
Zhu, X., Wang, Y., Dai, J., Yuan, L., Wei, Y.: Flow-guided feature aggregation for video object detection. In: Proceedings of the IEEE International Conference on Computer Vision (2017)
2017
Earlier work this paper cites.
Bertasius, G., Torresani, L., Shi, J.: Object detection in video with spatiotemporal sampling networks. In: European Conference on Computer Vision. pp. 331–346 (2018)
2018
Earlier work this paper cites.
Liu, S., Qi, L., Qin, H., Shi, J., Jia, J.: Path aggregation network for instance segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 8759–8768 (2018)
2018
Earlier work this paper cites.
Yang, L., Wang, Y., Xiong, X., Yang, J., Katsaggelos, A.K.: Efficient video object segmentation via network modulation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6499–6507 (2018)
2018
Earlier work this paper cites.
Bolya, D., Zhou, C., Xiao, F., Lee, Y.J.: Yolact: Real-time instance segmentation. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 9157–9166 (2019)
2019
Earlier work this paper cites.
Chen, K., Pang, J., Wang, J., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Shi, J., Ouyang, W., et al.: Hybrid task cascade for instance segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 4974–4983 (2019)
2019
Earlier work this paper cites.
Huang, Z., Huang, L., Gong, Y., Huang, C., Wang, X.: Mask scoring r-cnn. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6409–6418 (2019)
2019
Earlier work this paper cites.
Jiang, Z., Gao, P., Guo, C., Zhang, Q., Xiang, S., Pan, C.: Video object detection with locally-weighted deformable neighbors. In: Proceedings of the AAAI Conference on Artificial Intelligence (2019)
2019
Earlier work this paper cites.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al.: Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems 32
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Tian, Z., Shen, C., Chen, H., He, T.: Fcos: Fully convolutional one-stage object detection. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 9627–9636 (2019)
2019
Cited alongside, same era.
Voigtlaender, P., Chai, Y., Schroff, F., Adam, H., Leibe, B., Chen, L.C.: Feelvos: Fast end-to-end embedding learning for video object segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 9481–9490 (2019)
2019
Cited alongside, same era.
Wu, H., Chen, Y., Wang, N., Zhang, Z.: Sequence level semantics aggregation for video object detection. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 9217–9225 (2019)
2019
Cited alongside, same era.
Yang, L., Fan, Y., Xu, N.: Video instance segmentation. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 5188–5197 (2019)
2019
Cited alongside, same era.
Wang, Y., Xu, Z., Shen, H., Cheng, B., Yang, L.: Centermask: single shot instance segmentation with point representation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 9313–9321 (2020)
2020
Later among the works it cites.
Xie, E., Sun, P., Song, X., Wang, W., Liu, X., Liang, D., Shen, C., Luo, P.: Polarmask: Single shot instance segmentation with polar representation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 12193–12202 (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Fang, Y., Yang, S., Wang, X., Li, Y., Fang, C., Shan, Y., Feng, B., Liu, W.: Instances as queries. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 6910–6919 (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Athar, A., Mahadevan, S., Osep, A., Leal-Taixé, L., Leibe, B.: Stem-seg: Spatio-temporal embeddings for instance segmentation in videos. In: European Conference on Computer Vision. pp. 158–177 (2020)
2020
Cited alongside, same era.
Bertasius, G., Torresani, L.: Classifying, segmenting, and tracking object instances in video with mask propagation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 9739–9748 (2020)
2020
Cited alongside, same era.
Cao, J., Anwer, R.M., Cholakkal, H., Khan, F.S., Pang, Y., Shao, L.: Sipmask: Spatial information preservation for fast image and video instance segmentation. In: European Conference on Computer Vision. pp. 1–18 (2020)
2020
Cited alongside, same era.
Chen, H., Sun, K., Tian, Z., Shen, C., Huang, Y., Yan, Y.: Blendmask: Top-down meets bottom-up for instance segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 8573–8581 (2020)
2020
Cited alongside, same era.
Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for contrastive learning of visual representations. In: International Conference on Machine Learning. pp. 1597–1607 (2020)
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.: Momentum contrast for unsupervised visual representation learning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 9729–9738 (2020)
2020
Cited alongside, same era.
Hwang, S., Heo, M., Oh, S.W., Kim, S.J.: Video instance segmentation using inter-frame communication transformers. Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Ke, L., Li, X., Danelljan, M., Tai, Y.W., Tang, C.K., Yu, F.: Prototypical cross-attention networks for multiple object tracking and segmentation. Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Li, M., Li, S., Li, L., Zhang, L.: Spatial feature calibration and temporal fusion for effective one-stage video instance segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 11215–11224 (2021)
2021
Later among the works it cites.
Lin, H., Wu, R., Liu, S., Lu, J., Jia, J.: Video instance segmentation with a propose-reduce paradigm. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 1739–1748 (2021)
2021
Later among the works it cites.
Liu, D., Cui, Y., Tan, W., Chen, Y.: Sg-net: Spatial granularity network for one-stage video instance segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 9816–9825 (2021)
2021
Later among the works it cites.
Oksuz, K., Cam, B.C., Akbas, E., Kalkan, S.: Rank & sort loss for object detection and instance segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 3009–3018 (2021)
2021
Later among the works it cites.
Pang, J., Qiu, L., Li, X., Chen, H., Li, Q., Darrell, T., Yu, F.: Quasi-dense similarity learning for multiple object tracking. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 164–173 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Wang, T., Xu, N., Chen, K., Lin, W.: End-to-end video instance segmentation via spatial-temporal graph neural networks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 10797–10806 (2021)
2021
Later among the works it cites.
Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., Xia, H.: End-to-end video instance segmentation with transformers. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 8741–8750 (2021)
2021
Later among the works it cites.
Wu, J., Cao, J., Song, L., Wang, Y., Yang, M., Yuan, J.: Track to detect and segment: An online multi-object tracker. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 12352–12361 (2021)
2021
Later among the works it cites.
Xu, N., Yang, L., Yang, J., Yue, D., Fan, Y., Liang, Y., Huang, T.S.: Youtube-vis dataset 2021 version. https://youtube-vos.org/dataset/vis (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Ying, X., Li, X., Chuah, M.C.: Srnet: Spatial relation network for efficient single-stage instance segmentation in videos. In: Proceedings of the 29th ACM International Conference on Multimedia. pp. 347–356 (2021)
2021
Later among the works it cites.