Fetching the paper…
Reading the bibliography…
Can our video understanding systems perceive objects when a heavy occlusion exists in a scene? To answer this question, we collect a large-scale dataset called OVIS for occluded video instance segmentation, that is, to simultaneously detect, segment, and track instances in occluded scenes.
Perception 18
Nakayama, K., Shimojo, S., Silverman, G.H.: Stereoscopic depth: its relation to image segmentation, grouping, and the recognition of occluded objects · 1989
Earlier work this paper cites.
Pattern Recognition Letters 30
Brostow, G.J., Fauqueur, J., Cipolla, R.: Semantic object classes in video: A high-definition ground truth database · 2009
Earlier work this paper cites.
In: CVPR (2012)
Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? the kitti vision benchmark suite · 2012
Earlier work this paper cites.
IEEE TPAMI 36
Smeulders, A.W., Chu, D.M., Cucchiara, R., Calderara, S., Dehghan, A., Shah, M.: Visual tracking: An experimental survey · 2013
Earlier work this paper cites.
In: ECCV (2014)
Lin, T.Y., Maire, M., Belongie, S.J., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context · 2014
Earlier work this paper cites.
In: CVPR (2015)
Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Van Der Smagt, P., Cremers, D., Brox, T.: Flownet: Learning optical flow with convolutional networks · 2015
Earlier work this paper cites.
In: CVPR (2015)
Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation · 2015
Earlier work this paper cites.
IJCV 115
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al.: Imagenet large scale visual recognition challenge · 2015
Earlier work this paper cites.
arXiv preprint arXiv:1609.08675 (2016)
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., Vijayanarasimhan, S.: Youtube-8m: A large-scale video classification benchmark · 2016
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition (2016)
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B.: The cityscapes dataset for semantic urban scene understanding · 2016
Earlier work this paper cites.
In: ACCV (2016)
Fayyaz, M., Saffar, M.H., Sabokrou, M., Fathy, M., Klette, R., Huang, F.: Stfcn: spatio-temporal fcn for semantic video segmentation · 2016
Earlier work this paper cites.
In: CVPR (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Earlier work this paper cites.
arXiv preprint arXiv:1603.00831 (2016)
Milan, A., Leal-Taixé, L., Reid, I., Roth, S., Schindler, K.: Mot16: A benchmark for multi-object tracking · 2016
Earlier work this paper cites.
In: CVPR (2016)
Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., Sorkine-Hornung, A.: A benchmark dataset and evaluation methodology for video object segmentation · 2016
Earlier work this paper cites.
IEEE TPAMI 40
Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs · 2017
Earlier work this paper cites.
In: ICCV, pp. 4836–4845 (2017)
Chu, Q., Ouyang, W., Li, H., Wang, X., Liu, B., Yu, N.: Online multi-object tracking using cnn-based single object tracker with spatial-temporal attention mechanism · 2017
Earlier work this paper cites.
In: ICCV (2017)
Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., Wei, Y.: Deformable convolutional networks · 2017
Earlier work this paper cites.
arXiv preprint arXiv:1708.04552 (2017)
DeVries, T., Taylor, G.W.: Improved regularization of convolutional neural networks with cutout · 2017
Earlier work this paper cites.
In: ICCV, pp. 1301–1310 (2017)
Dwibedi, D., Misra, I., Hebert, M.: Cut, paste and learn: Surprisingly easy synthesis for instance detection · 2017
Earlier work this paper cites.
In: CVPR (2017)
He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask r-cnn · 2017
Earlier work this paper cites.
In: CVPR, pp. 4507–4515 (2017)
Hosang, J., Benenson, R., Schiele, B.: Learning non-maximum suppression · 2017
Earlier work this paper cites.
In: CVPR (2017)
Khoreva, A., Perazzi, F., Benenson, R., Schiele, B., Sorkine-Hornung, A.: Learning video object segmentation from static images · 2017
Earlier work this paper cites.
In: CVPR (2017)
Tokmakov, P., Alahari, K., Schmid, C.: Learning motion patterns in videos · 2017
Earlier work this paper cites.
In: BMVC (2017)
Voigtlaender, P., Leibe, B.: Online adaptation of convolutional neural networks for video object segmentation · 2017
Earlier work this paper cites.
In: CVPR (2017)
Xie, S., Girshick, R., Dollár, P., Tu, Z., He, K.: Aggregated residual transformations for deep neural networks · 2017
Earlier work this paper cites.
In: CVPR, pp. 3213–3221 (2017)
Zhang, S., Benenson, R., Schiele, B.: Citypersons: A diverse dataset for pedestrian detection · 2017
Earlier work this paper cites.
In: CVPR (2017)
Zhu, X., Xiong, Y., Dai, J., Yuan, L., Wei, Y.: Deep feature flow for video recognition · 2017
Earlier work this paper cites.
In: ECCV, pp. 331–346 (2018)
Bertasius, G., Torresani, L., Shi, J.: Object detection in video with spatiotemporal sampling networks · 2018
Earlier work this paper cites.
In: ECCV (2018)
Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., Adam, H.: Encoder-decoder with atrous separable convolution for semantic image segmentation · 2018
Earlier work this paper cites.
In: ECCV (2018)
Hu, Y.T., Huang, J.B., Schwing, A.G.: Videomatch: Matching based video object segmentation · 2018
Earlier work this paper cites.
In: CVPR (2018)
Li, S., Seybold, B., Vorobyov, A., Fathi, A., Kuo, C.C.J.: Instance embedding transfer to unsupervised video object segmentation · 2018
Earlier work this paper cites.
In: ECCV (2018)
Li, X., Loy, C.C.: Video object segmentation with joint re-identification and attention-aware mask propagation · 2018
Earlier work this paper cites.
In: CVPR (2018)
Nilsson, D., Sminchisescu, C.: Semantic video segmentation by gated recurrent flow propagation · 2018
Cited alongside, same era.
In: CVPR (2018)
Oh, S.W., Lee, J.Y., Sunkavalli, K., Kim, S.J.: Fast video object segmentation by reference-guided mask propagation · 2018
Cited alongside, same era.
arXiv preprint arXiv:1805.00123 (2018)
Shao, S., Zhao, Z., Li, B., Xiao, T., Yu, G., Zhang, X., Sun, J.: Crowdhuman: A benchmark for detecting human in a crowd · 2018
Cited alongside, same era.
In: CVPR, pp. 7794–7803 (2018)
Wang, X., Girshick, R., Gupta, A., He, K.: Non-local neural networks · 2018
Cited alongside, same era.
In: CVPR (2018)
Wang, X., Xiao, T., Jiang, Y., Shao, S., Sun, J., Shen, C.: Repulsion loss: Detecting pedestrians in a crowd · 2018
Cited alongside, same era.
In: ECCV (2018)
Xu, N., Yang, L., Fan, Y., Yang, J., Yue, D., Liang, Y., Price, B., Cohen, S., Huang, T.: Youtube-vos: Sequence-to-sequence video object segmentation · 2018
In: CVPR (2020)
Chu, X., Zheng, A., Zhang, X., Sun, J.: Detection in crowded scenes: One proposal, multiple predictions · 2020
Later among the works it cites.
In: ECCV, pp. 715–733. Springer (2020)
Devaranjan, J., Kar, A., Fidler, S.: Meta-sim2: Unsupervised learning of scene structure for synthetic data generation · 2020
Later among the works it cites.
In: CVPR (2020)
Kim, D., Woo, S., Lee, J.Y., Kweon, I.S.: Video panoptic segmentation · 2020
Later among the works it cites.
In: CVPR (2020)
Kirillov, A., Wu, Y., He, K., Girshick, R.: Pointrend: Image segmentation as rendering · 2020
Later among the works it cites.
In: CVPR, pp. 8940–8949 (2020)
Kortylewski, A., He, J., Liu, Q., Yuille, A.L.: Compositional convolutional neural networks: A deep architecture with innate robustness to partial occlusion · 2020
Later among the works it cites.
In: WACV, pp. 1333–1341 (2020)
Kortylewski, A., Liu, Q., Wang, H., Zhang, Z., Yuille, A.: Combining compositional models and deep networks for robust object classification under occlusion · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In: ECCV (2018)
Zhang, S., Wen, L., Bian, X., Lei, Z., Li, S.Z.: Occlusion-aware r-cnn: Detecting pedestrians in a crowd · 2018
Cited alongside, same era.
In: ECCV, pp. 135–151 (2018)
Zhou, C., Yuan, J.: Bi-box regression for pedestrian detection and occlusion estimation · 2018
Cited alongside, same era.
In: ECCV, pp. 366–382 (2018)
Zhu, J., Yang, H., Liu, N., Kim, M., Zhang, W., Yang, M.H.: Online multi-object tracking with dual matching attention networks · 2018
Cited alongside, same era.
arXiv (2019)
Caelles, S., Pont-Tuset, J., Perazzi, F., Montes, A., Maninis, K.K., Van Gool, L.: The 2019 davis challenge on vos: Unsupervised multi-object segmentation · 2019
Cited alongside, same era.
In: CVPR (2019)
Gupta, A., Dollar, P., Girshick, R.: Lvis: A dataset for large vocabulary instance segmentation · 2019
Cited alongside, same era.
In: CVPR (2019)
Huang, Z., Huang, L., Gong, Y., Huang, C., Wang, X.: Mask scoring r-cnn · 2019
Cited alongside, same era.
In: CVPR, pp. 10720–10729 (2020)
Lazarow, J., Lee, K., Shi, K., Tu, Z.: Learning instance occlusion for panoptic segmentation · 2020
Later among the works it cites.
In: CVPR (2020)
Li, Q., Qi, X., Torr, P.H.: Unifying training and inference for panoptic segmentation · 2020
Later among the works it cites.
NeurIPS 33
Li, Y., Xu, N., Peng, J., See, J., Lin, W.: Delving into the cyclic mechanism in semi-supervised video object segmentation · 2020
Later among the works it cites.
In: CVPR (2020)
Lin, C.C., Hung, Y., Feris, R., He, L.: Video instance segmentation tracking with a modified vae architecture · 2020
Later among the works it cites.
In: IJCAI, pp. 530–536 (2020)
Liu, Q., Chu, Q., Liu, B., Yu, N.: Gsm: Graph similarity model for multi-object tracking · 2020
Later among the works it cites.
arXiv preprint arXiv:2011.14503 (2020)
Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., Xia, H.: End-to-end video instance segmentation with transformers · 2020
Later among the works it cites.
Computer Vision and Image Understanding 193
Wen, L., Du, D., Cai, Z., Lei, Z., Chang, M.C., Qi, H., Lim, J., Yang, M.H., Lyu, S.: Ua-detrac: A new benchmark and protocol for multi-object detection and tracking · 2020
Later among the works it cites.
In: ACM Multimedia (2020)
Wu, J., Song, L., Wang, T., Zhang, Q., Yuan, J.: Forest r-cnn: Large-vocabulary long-tailed object detection and instance segmentation · 2020
Later among the works it cites.
In: CVPR (2020)
Wu, J., Zhou, C., Yang, M., Zhang, Q., Li, Y., Yuan, J.: Temporal-context enhanced detection of heavily occluded pedestrians · 2020
Later among the works it cites.
In: ECCV (2020)
Xu, Z., Zhang, W., Tan, X., Yang, W., Huang, H., Wen, S., Ding, E., Huang, L.: Segment as points for efficient online multi-object tracking and segmentation · 2020
Later among the works it cites.
In: CVPR (2020)
Zhan, X., Pan, X., Dai, B., Liu, Z., Lin, D., Loy, C.C.: Self-supervised scene de-occlusion · 2020
Later among the works it cites.
In: ICCV (2021)
Fang, Y., Yang, S., Wang, X., Li, Y., Fang, C., Shan, Y., Feng, B., Liu, W.: Instances as queries · 2021
Closest in time.
In: CVPR, pp. 2918–2928 (2021)
Ghiasi, G., Cui, Y., Srinivas, A., Qian, R., Lin, T.Y., Cubuk, E.D., Le, Q.V., Zoph, B.: Simple copy-paste is a strong data augmentation method for instance segmentation · 2021
Closest in time.
arXiv preprint arXiv:2111.06377 (2021)
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R.: Masked autoencoders are scalable vision learners · 2021
Closest in time.
arXiv preprint arXiv:2106.03299 (2021)
Hwang, S., Heo, M., Oh, S.W., Kim, S.J.: Video instance segmentation using inter-frame communication transformers · 2021
Closest in time.
In: CVPR (2021)
Ke, L., Tai, Y.W., Tang, C.K.: Deep occlusion-aware instance segmentation with overlapping bilayers · 2021
Closest in time.
IJCV 129
Kortylewski, A., Liu, Q., Wang, A., Sun, Y., Yuille, A.: Compositional convolutional neural networks: A robust and interpretable model for object recognition under occlusion · 2021
Closest in time.
In: CVPR (2021)
Li, M., Li, S., Li, L., Zhang, L.: Spatial feature calibration and temporal fusion for effective one-stage video instance segmentation · 2021
Closest in time.
In: CVPR (2021)
Liu, D., Cui, Y., Tan, W., Chen, Y.: Sg-net: Spatial granularity network for one-stage video instance segmentation · 2021
Closest in time.
In: ICCV, pp. 10012–10022 (2021)
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows · 2021
Closest in time.
In: Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2021)
Qi, J., Gao, Y., Hu, Y., Wang, X., Liu, X., Bai, X., Belongie, S., Yuille, A., Torr, P., Bai, S.: Occluded video instance segmentation: Dataset and iccv 2021 challenge · 2021
Closest in time.
In: CVPR (2021)
Wang, H., Jiang, X., Ren, H., Hu, Y., Bai, S.: Swiftnet: Real-time video object segmentation · 2021
Closest in time.
arXiv preprint arXiv:2104.04691 (2021)
Wang, W., Feiszli, M., Wang, H., Tran, D.: Unidentified video objects: A benchmark for dense, open-world segmentation · 2021
Closest in time.
In: CVPR (2021)
Wu, J., Cao, J., Song, L., Wang, Y., Yang, M., Yuan, J.: Track to detect and segment: An online multi-object tracker · 2021
Closest in time.
In: ICCV (2021)
Yang, S., Fang, Y., Wang, X., Li, Y., Fang, C., Shan, Y., Feng, B., Liu, W.: Crossover learning for fast online video instance segmentation · 2021
Closest in time.