Fetching the paper…
Reading the bibliography…
Open-vocabulary Video Instance Segmentation (OpenVIS) can simultaneously detect, segment, and track arbitrary object categories in a video, without being constrained to categories seen during training.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Simple online and realtime tracking
Bewley, A.; Ge, Z.; Ott, L.; Ramos, F.; and Upcroft, B. 2016 · 2016
Earlier work this paper cites.
Mask r-cnn
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
Zero-shot semantic segmentation
Bucher, M.; Vu, T.-H.; Cord, M.; and Pérez, P. 2019 · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Gupta, A.; Dollar, P.; and Girshick, R. 2019 · 2019
Earlier work this paper cites.
Video instance segmentation
Yang, L.; Fan, Y.; and Xu, N. 2019 · 2019
Earlier work this paper cites.
Classifying, segmenting, and tracking object instances in video with mask propagation
Bertasius, G.; and Torresani, L. 2020 · 2020
Earlier work this paper cites.
Sipmask: Spatial information preservation for fast image and video instance segmentation
Cao, J.; Anwer, R. M.; Cholakkal, H.; Khan, F. S.; Pang, Y.; and Shao, L. 2020 · 2020
Earlier work this paper cites.
Is Space-Time Attention All You Need for Video Understanding?
Bertasius, G.; Wang, H.; and Torresani, L. 2021 · 2021
Earlier work this paper cites.
Mask2former for video instance segmentation
Cheng, B.; Choudhuri, A.; Misra, I.; Kirillov, A.; Girdhar, R.; and Schwing, A. G. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Image retrieval on real-life images with pre-trained vision-and-language models
Liu, Z.; Rodriguez-Opazo, C.; Teney, D.; and Gould, S. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Seqformer: a frustratingly simple model for video instance segmentation
Wu, J.; Jiang, Y.; Zhang, W.; Bai, X.; and Bai, S. 2021 · 2021
Cited alongside, same era.
A simple baseline for zero-shot semantic segmentation with pre-trained vision-language model
Xu, M.; Zhang, Z.; Wei, F.; Lin, Y.; Cao, Y.; Hu, H.; and Bai, X. 2021 · 2021
Cited alongside, same era.
Crossover learning for fast online video instance segmentation
Yang, S.; Fang, Y.; Wang, X.; Li, Y.; Fang, C.; Shan, Y.; Feng, B.; and Liu, W. 2021 · 2021
Cited alongside, same era.
FILIP: fine-grained interactive language-image pre-training
Yao, L.; Huang, R.; Hou, L.; Lu, G.; Niu, M.; Xu, H.; Liang, X.; Li, Z.; Jiang, X.; and Xu, C. 2021 · 2021
Cited alongside, same era.
Masked-attention mask transformer for universal image segmentation
Cheng, B.; Misra, I.; Schwing, A. G.; Kirillov, A.; and Girdhar, R. 2022 · 2022
DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in Videos
Wang, Y.; Sun, Y.; Song, W.; Gao, S.; Huang, Y.; Chen, Z.; Ge, W.; and Zhang, W. 2022 · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision
Zhou, X.; Girdhar, R.; Joulin, A.; Krähenbühl, P.; and Misra, I. 2022 · 2022
Later among the works it cites.
BURST: A Benchmark for Unifying Object Recognition, Segmentation and Tracking in Video
Athar, A.; Luiten, J.; Voigtlaender, P.; Khurana, T.; Dave, A.; Leibe, B.; and Ramanan, D. 2023 · 2023
Closest in time.
Vision Transformers Need Registers
Darcet, T.; Oquab, M.; Mairal, J.; and Bojanowski, P. 2023 · 2023
Closest in time.
Open-Vocabulary Semantic Segmentation with Decoupled One-Pass Network
Han, C.; Zhong, Y.; Li, D.; Han, K.; and Ma, L. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Decoupling zero-shot semantic segmentation
Ding, J.; Xue, N.; Xia, G.-S.; and Dai, D. 2022 · 2022
Cited alongside, same era.
Adaptive online mutual learning bi-decoders for video object segmentation
Guo, P.; Zhang, W.; Li, X.; and Zhang, W. 2022 · 2022
Cited alongside, same era.
Visolo: Grid-based space-time aggregation for efficient online video instance segmentation
Han, S. H.; Hwang, S.; Oh, S. W.; Park, Y.; Kim, H.; Kim, M.-J.; and Kim, S. J. 2022 · 2022
Cited alongside, same era.
LVOS: A Benchmark for Long-term Video Object Segmentation
Hong, L.; Chen, W.; Liu, Z.; Zhang, W.; Guo, P.; Chen, Z.; and Zhang, W. 2022 · 2022
Cited alongside, same era.
Scaling up vision-language pre-training for image captioning
Hu, X.; Gan, Z.; Wang, J.; Yang, Z.; Liu, Z.; Lu, Y.; and Wang, L. 2022 · 2022
Cited alongside, same era.
Minvis: A minimal video instance segmentation framework without video-based training
Huang, D.-A.; Yu, Z. x.; xxxx, x.; and xxxxxxx, x. 2022 · 2022
Cited alongside, same era.
Unsupervised prompt learning for vision-language models
Huang, T.; Chu, J.; and Wei, F. 2022 · 2022
Cited alongside, same era.
A generalized framework for video instance segmentation
Heo, M.; Hwang, S.; Hyun, J.; Kim, H.; Oh, S. W.; Lee, J.-Y.; and Kim, S. J. 2023 · 2023
Closest in time.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023 · 2023
Closest in time.
Videochat: Chat-centric video understanding
Li, K.; He, Y.; Wang, Y.; Li, Y.; Wang, W.; Luo, P.; Wang, Y.; Wang, L.; and Qiao, Y. 2023 · 2023
Closest in time.
Segment Anything Meets Point Tracking
Rajič, F.; Ke, L.; Tai, Y.-W.; Tang, C.-K.; Danelljan, M.; and Yu, F. 2023 · 2023
Closest in time.
Towards open-vocabulary video instance segmentation
Wang, H.; Yan, C.; Wang, S.; Jiang, X.; Tang, X.; Hu, Y.; Xie, W.; and Gavves, E. 2023 · 2023
Closest in time.
Side adapter network for open-vocabulary semantic segmentation
Xu, M.; Zhang, Z.; Wei, F.; Hu, H.; and Bai, X. 2023 · 2023
Closest in time.
Track anything: Segment anything meets videos
Yang, J.; Gao, M.; Li, Z.; Gao, S.; Wang, F.; and Zheng, F. 2023 · 2023
Closest in time.
DVIS: Decoupled Video Instance Segmentation Framework
Zhang, T.; Tian, X.; Wu, Y.; Ji, S.; Wang, X.; Zhang, Y.; and Wan, P. 2023 · 2023
Closest in time.
Memory Network with Pixel-level Spatio-Temporal Learning for Visual Object Tracking
Zhou, Z.; Zhou, X.; Chen, Z.; Guo, P.; Liu, Q.-Y.; and Zhang, W. 2023 · 2023
Closest in time.