Fetching the paper…
Reading the bibliography…
We propose the new task 'open-world video instance segmentation and captioning'.
Algorithms for the assignment and transportation problems
James Munkres · 1957
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Youtube-vos: Sequence-to-sequence video object segmentation
Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang · 2018
Earlier work this paper cites.
Videomatch: Matching based video object segmentation
Yuan-Ting Hu, Jia-Bin Huang, and Alexander G Schwing · 2018
Earlier work this paper cites.
Fast online object tracking and segmentation: A unifying approach
Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip HS Torr · 2019
Earlier work this paper cites.
Fast video object segmentation via dynamic targeting network
Lu Zhang, Zhe Lin, Jianming Zhang, Huchuan Lu, and You He · 2019
Earlier work this paper cites.
Video object segmentation using space-time memory networks
Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim · 2019
Earlier work this paper cites.
Video instance segmentation
Linjie Yang, Yuchen Fan, and Ning Xu · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Earlier work this paper cites.
Where does it exist: Spatio-temporal video grounding for multi-form sentences
Zhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang, Huasheng Liu, and Lianli Gao · 2020
Earlier work this paper cites.
Classifying, segmenting, and tracking object instances in video with mask propagation
Gedas Bertasius and Lorenzo Torresani · 2020
Earlier work this paper cites.
Stem-seg: Spatio-temporal embeddings for instance segmentation in videos
Ali Athar, Sabarinath Mahadevan, Aljosa Osep, Laura Leal-Taixé, and Bastian Leibe · 2020
Earlier work this paper cites.
Learning a neural solver for multiple object tracking
Guillem Braso and Laura Leal-Taixe · 2020
Earlier work this paper cites.
Lifted disjoint paths with application in multiple object tracking
Andrea Hornakova, Roberto Henschel, Bodo Rosenhahn, and Paul Swoboda · 2020
Earlier work this paper cites.
Track to reconstruct and reconstruct to track
Jonathon Luiten, Tobias Fischer, and Bastian Leibe · 2020
Earlier work this paper cites.
Segment as points for efficient online multi-object tracking and segmentation
Zhenbo Xu, Wei Zhang, Xiao Tan, Wei Yang, Huan Huang, Shilei Wen, Errui Ding, and Liusheng Huang · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Fourier features let networks learn high frequency functions in low dimensional domains
Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng · 2020
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Cited alongside, same era.
Occluded video instance segmentation
Jiyang Qi, Yan Gao, Yao Hu, Xinggang Wang, Xiaoyu Liu, Xiang Bai, Serge Belongie, Alan Yuille, Philip HS Torr, and Song Bai · 2021
Cited alongside, same era.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Cited alongside, same era.
Crossover learning for fast online video instance segmentation
Shusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu · 2021
Video mask transfiner for high-quality video instance segmentation
Lei Ke, Henghui Ding, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, and Fisher Yu · 2022
Later among the works it cites.
Instanceformer: An online video instance segmentation framework
Rajat Koner, Tanveer Hannan, Suprosanna Shit, Sahand Sharifzadeh, Matthias Schubert, Thomas Seidl, and Volker Tresp · 2022
Later among the works it cites.
Vita: Video instance segmentation via object token association
Miran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee, and Seon Joo Kim · 2022
Later among the works it cites.
Video instance segmentation in an open-world
Omkar Thawakar, Sanath Narayan, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Jorma Laaksonen, Mubarak Shah, and Fahad Shahbaz Khan · 2023
Later among the works it cites.
Burst: A benchmark for unifying object recognition, segmentation and tracking in video
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Assignment-space-based multi-object tracking and segmentation
Anwesa Choudhuri, Girish Chowdhary, and Alexander G. Schwing · 2021
Cited alongside, same era.
Seqformer: a frustratingly simple model for video instance segmentation
Junfeng Wu, Yi Jiang, Wenqing Zhang, Xiang Bai, and Song Bai · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Cited alongside, same era.
Minvis: A minimal video instance segmentation framework without video-based training
De-An Huang, Zhiding Yu, and Anima Anandkumar · 2022
Cited alongside, same era.
Image segmentation using text and image prompts
Timo Lüddecke and Alexander Ecker · 2022
Cited alongside, same era.
Tubeformer-deeplab: Video mask transformer
Dahun Kim, Jun Xie, Huiyu Wang, Siyuan Qiao, Qihang Yu, Hong-Seok Kim, Hartwig Adam, In So Kweon, and Liang-Chieh Chen · 2022
Cited alongside, same era.
Ali Athar, Jonathon Luiten, Paul Voigtlaender, Tarasha Khurana, Achal Dave, Bastian Leibe, and Deva Ramanan · 2023
Later among the works it cites.
Dense video object captioning from disjoint supervision
Xingyi Zhou, Anurag Arnab, Chen Sun, and Cordelia Schmid · 2023
Later among the works it cites.
Tracking anything with decoupled video segmentation
Ho Kei Cheng, Seoung Wug Oh, Brian Price, Alexander Schwing, and Joon-Young Lee · 2023
Later among the works it cites.
Capdet: Unifying dense captioning and open-world detection pretraining
Yanxin Long, Youpeng Wen, Jianhua Han, Hang Xu, Pengzhen Ren, Wei Zhang, Shen Zhao, and Xiaodan Liang · 2023
Later among the works it cites.
Honeybee: Locality-enhanced projector for multimodal llm
Junbum Cha, Wooyoung Kang, Jonghwan Mun, and Byungseok Roh · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Later among the works it cites.
Context-aware relative object queries to unify video instance and panoptic segmentation
Anwesa Choudhuri, Girish Chowdhary, and Alexander G. Schwing · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Later among the works it cites.
Generalized decoding for pixel, image, and language
Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, et al · 2023
Later among the works it cites.
Oneformer: One transformer to rule universal image segmentation
Jitesh Jain, Jiachen Li, Mang Tik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi · 2023
Later among the works it cites.