Fetching the paper…
Reading the bibliography…
We propose a novel end-to-end solution for video instance segmentation (VIS) based on transformers.
The hungarian method for the assignment problem
Kuhn, H. W · 1955
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., M. Maire, S. Belongie, et al · 2014
Earlier work this paper cites.
End-to-end people detection in crowded scenes
Stewart, R., M. Andriluka, A. Y. Ng · 2016
Earlier work this paper cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Milletari, F., N. Navab, S.-A. Ahmadi · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., X. Zhang, S. Ren, et al · 2016
Earlier work this paper cites.
Mask r-cnn
He, K., G. Gkioxari, P. Dollar, et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., N. Shazeer, N. Parmar, et al · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T.-Y., P. Dollar, R. Girshick, et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T.-Y., P. Goyal, R. Girshick, et al · 2017
Earlier work this paper cites.
Deformable convolutional networks
Dai, J., H. Qi, Y. Xiong, et al · 2017
Earlier work this paper cites.
Non-local neural networks
Wang, X., R. Girshick, A. Gupta, et al · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Tran, D., H. Wang, L. Torresani, et al · 2018
Earlier work this paper cites.
Cascade r-cnn: Delving into high quality object detection
Cai, Z., N. Vasconcelos · 2018
Cited alongside, same era.
Video instance segmentation
Yang, L., Y. Fan, N. Xu · 2019
Cited alongside, same era.
Yolact: Real-time instance segmentation
Bolya, D., C. Zhou, F. Xiao, et al · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., M.-W. Chang, K. Lee, et al · 2019
Cited alongside, same era.
Panoptic segmentation
Kirillov, A., K. He, R. Girshick, et al · 2019
Cited alongside, same era.
Detectron2
Wu, Y., A. Kirillov, F. Massa, et al · 2019
Cited alongside, same era.
End-to-end video instance segmentation with transformers
Wang, Y., Z. Xu, X. Wang, et al · 2020
Later among the works it cites.
End-to-end object detection with transformers
Carion, N., F. Massa, G. Synnaeve, et al · 2020
Later among the works it cites.
Axial-deeplab: Stand-alone axial-attention for panoptic segmentation
Wang, H., Y. Zhu, B. Green, et al · 2020
Later among the works it cites.
Stem-seg: Spatio-temporal embeddings for instance segmentation in videos
Athar, A., S. Mahadevan, A. Ošep, et al · 2020
Later among the works it cites.
Compfeat: Comprehensive feature aggregation for video instance segmentation
Fu, Y., L. Yang, D. Liu, et al · 2020
Later among the works it cites.
Crossover learning for fast online video instance segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decoupled weight decay regularization
Loshchilov, I., F. Hutter · 2019
Cited alongside, same era.
Sipmask: Spatial information preservation for fast image and video instance segmentation
Cao, J., R. M. Anwer, H. Cholakkal, et al · 2020
Cited alongside, same era.
Conditional convolutions for instance segmentation
Tian, Z., C. Shen, H. Chen · 2020
Cited alongside, same era.
Blendmask: Top-down meets bottom-up for instance segmentation
Chen, H., K. Sun, Z. Tian, et al · 2020
Cited alongside, same era.
Multiple object tracking: A literature review
Luo, W., J. Xing, A. Milan, et al · 2020
Cited alongside, same era.
Classifying, segmenting, and tracking object instances in video with mask propagation
Bertasius, G., L. Torresani · 2020
Cited alongside, same era.
Yang, S., Y. Fang, X. Wang, et al · 2021
Closest in time.
Sg-net: Spatial granularity network for one-stage video instance segmentation
Liu, D., Y. Cui, W. Tan, et al · 2021
Closest in time.
Video instance segmentation with a propose-reduce paradigm
Lin, H., R. Wu, S. Liu, et al · 2021
Closest in time.
Is space-time attention all you need for video understanding?
Bertasius, G., H. Wang, L. Torresani · 2021
Closest in time.
Max-deeplab: End-to-end panoptic segmentation with mask transformers
Wang, H., Y. Zhu, H. Adam, et al · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., L. Beyer, A. Kolesnikov, et al · 2021
Closest in time.
Vision transformers for dense prediction
Ranftl, R., A. Bochkovskiy, V. Koltun · 2021
Closest in time.