Fetching the paper…
Reading the bibliography…
In this paper, we propose an Audio-Language-Referenced SAM 2 (AL-Ref-SAM 2) pipeline to explore the training-free paradigm for audio and language-referenced video object segmentation, namely AVS and RVOS tasks.
The 2017 davis challenge on video object segmentation
Pont-Tuset, J.; Perazzi, F.; Caelles, S.; Arbeláez, P.; Sorkine-Hornung, A.; and Van Gool, L. 2017 · 2017
Earlier work this paper cites.
Youtube-vos: Sequence-to-sequence video object segmentation
Xu, N.; Yang, L.; Fan, Y.; Yang, J.; Yue, D.; Liang, Y.; Price, B.; Cohen, S.; and Huang, T. 2018 · 2018
Earlier work this paper cites.
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
Zhu, Z.; Feng, X.; Chen, D.; Yuan, J.; Qiao, C.; and Hua, G. 2024 · 2018
Earlier work this paper cites.
Video object segmentation with language referring expressions
Khoreva, A.; Rohrbach, A.; and Schiele, B. 2019 · 2019
Earlier work this paper cites.
Referring image segmentation via cross-modal progressive comprehension
Huang, S.; Hui, T.; Liu, S.; Li, G.; Wei, Y.; Han, J.; Liu, L.; and Li, B. 2020 · 2020
Earlier work this paper cites.
Linguistic structure guided context modeling for referring image segmentation
Hui, T.; Liu, S.; Huang, S.; Li, G.; Yu, S.; Zhang, F.; and Han, J. 2020 · 2020
Earlier work this paper cites.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seo, S.; Lee, J.-Y.; and Han, B. 2020 · 2020
Earlier work this paper cites.
Collaborative spatial-temporal modeling for language-queried video actor segmentation
Hui, T.; Huang, S.; Liu, S.; Ding, Z.; Li, G.; Wang, W.; Han, J.; and Wang, F. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Human-centric spatio-temporal video grounding with visual transformers
Tang, Z.; Liao, Y.; Liu, S.; Li, G.; Jin, X.; Jiang, H.; Yu, Q.; and Xu, D. 2021 · 2021
Earlier work this paper cites.
End-to-end referring video object segmentation with multimodal transformers
Botach, A.; Zheltonozhskii, E.; and Baskin, C. 2022 · 2022
Earlier work this paper cites.
Beats: Audio pre-training with acoustic tokenizers
Chen, S.; Wu, Y.; Wang, C.; Liu, S.; Tompkins, D.; Chen, Z.; and Wei, F. 2022 · 2022
Earlier work this paper cites.
Language-bridged spatial-temporal interaction for referring video object segmentation
Ding, Z.; Hui, T.; Huang, J.; Wei, X.; Han, J.; and Liu, S. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Language as queries for referring video object segmentation
Wu, J.; Jiang, Y.; Sun, P.; Yuan, Z.; and Luo, P. 2022 · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Earlier work this paper cites.
Audio–visual segmentation
Zhou, J.; Wang, J.; Zhang, J.; Sun, W.; Zhang, J.; Birchfield, S.; Guo, D.; Kong, L.; Wang, M.; and Zhong, Y. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Cheng, Y.; Li, L.; Xu, Y.; Li, X.; Yang, Z.; Wang, W.; and Yang, Y. 2023 · 2023
Cited alongside, same era.
MeViS: A large-scale benchmark for video segmentation with motion expressions
Ding, H.; Liu, C.; He, S.; Jiang, X.; and Loy, C. C. 2023 · 2023
Cited alongside, same era.
Html: Hybrid temporal-scale multimodal learning framework for referring video object segmentation
Han, M.; Wang, Y.; Li, Z.; Yao, L.; Chang, X.; and Qiao, Y. 2023 · 2023
Cited alongside, same era.
Audio-visual segmentation with semantics
Zhou, J.; Shen, X.; Wang, J.; Zhang, J.; Sun, W.; Zhang, J.; Birchfield, S.; Guo, D.; Kong, L.; Wang, M.; et al. 2023 · 2023
Later among the works it cites.
Zhu, B.; Lin, B.; Ning, M.; Yan, Y.; Cui, J.; Wang, H.; Pang, Y.; Jiang, W.; Zhang, J.; Li, Z.; et al. 2023 · 2023
Later among the works it cites.
Unsupervised Audio-Visual Segmentation with Modality Alignment
Bhosale, S.; Yang, H.; Kanojia, D.; Deng, J.; and Zhu, X. 2024 · 2024
Closest in time.
Unraveling Instance Associations: A Closer Look for Audio-Visual Segmentation
Chen, Y.; Liu, Y.; Wang, H.; Liu, F.; Wang, C.; Frazer, H.; and Carneiro, G. 2024 · 2024
Closest in time.
Avsegformer: Audio-visual segmentation with transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Discovering sounding objects by audio queries for audio visual segmentation
Huang, S.; Li, H.; Wang, Y.; Zhu, H.; Dai, J.; Han, J.; Rong, W.; and Liu, S. 2023 · 2023
Cited alongside, same era.
Language-aware spatial-temporal collaboration for referring video segmentation
Hui, T.; Liu, S.; Ding, Z.; Huang, S.; Li, G.; Wang, W.; Liu, L.; and Han, J. 2023 · 2023
Cited alongside, same era.
Segment anything
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023 · 2023
Cited alongside, same era.
Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation
Li, K.; Yang, Z.; Chen, L.; Yang, Y.; and Xiao, J. 2023 · 2023
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Li, C.; Yang, J.; Su, H.; Zhu, J.; et al. 2023 · 2023
Cited alongside, same era.
Spectrum-guided multi-granularity referring video object segmentation
Miao, B.; Bennamoun, M.; Gao, Y.; and Mian, A. 2023 · 2023
Cited alongside, same era.
Temporal collection and distribution for referring video object segmentation
Tang, J.; Zheng, G.; and Yang, S. 2023 · 2023
Cited alongside, same era.
Gao, S.; Chen, Z.; Chen, G.; Wang, W.; and Lu, T. 2024 · 2024
Closest in time.
Decoupling static and hierarchical motion perception for referring video segmentation
He, S.; and Ding, H. 2024 · 2024
Closest in time.
GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation
Lin, C.-S.; Liu, I.; Chen, M.-H.; Wang, C.-Y.; Liu, S.; Wang, Y.-C. F.; et al. 2024 · 2024
Closest in time.
Primitivenet: decomposing the global constraints for referring segmentation
Liu, C.; Jiang, X.; and Ding, H. 2024 · 2024
Closest in time.
Soc: Semantic-assisted object cluster for referring video object segmentation
Luo, Z.; Xiao, Y.; Liu, Y.; Li, S.; Wang, Y.; Tang, Y.; Li, X.; and Yang, Y. 2024 · 2024
Closest in time.
Weakly-supervised audio-visual segmentation
Mo, S.; and Raj, B. 2024 · 2024
Closest in time.
Sam 2: Segment anything in images and videos
Ravi, N.; Gabeur, V.; Hu, Y.-T.; Hu, R.; Ryali, C.; Ma, T.; Khedr, H.; Rädle, R.; Rolland, C.; Gustafson, L.; et al. 2024 · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks
Ren, T.; Liu, S.; Zeng, A.; Lin, J.; Li, K.; Cao, H.; Chen, J.; Huang, X.; Chen, Y.; Yan, F.; et al. 2024 · 2024
Closest in time.
Mask-enhanced segment anything model for tumor lesion semantic segmentation
Shi, H.; Han, S.; Huang, S.; Liao, Y.; Li, G.; Kong, X.; Zhu, H.; Wang, X.; and Liu, S. 2024 · 2024
Closest in time.
Prompting segmentation with sound is generalizable audio-visual source localizer
Wang, Y.; Liu, W.; Li, G.; Ding, J.; Hu, D.; and Li, X. 2024 · 2024
Closest in time.
Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation
Yang, Q.; Nie, X.; Li, T.; Gao, P.; Guo, Y.; Zhen, C.; Yan, P.; and Xiang, S. 2024 · 2024
Closest in time.
Losh: Long-short text joint prediction network for referring video object segmentation
Yuan, L.; Shi, M.; Yue, Z.; and Chen, Q. 2024 · 2024
Closest in time.