Fetching the paper…
Reading the bibliography…
Audio-Visual Segmentation (AVS) aims to extract the sounding object from a video frame, which is represented by a pixel-wise segmentation mask for application scenarios such as multi-modal video editing, augmented reality, and intelligent robot systems.
X. Wu, H. Gong, P. Chen, Z. Zhong, and Y. Xu, “Surveillance robot utilizing video and audio information,” Journal of Intelligent and Robotic Systems , vol. 55, pp. 403–421, 2009
2009
Earlier work this paper cites.
T. Mensink, J. Verbeek, F. Perronnin, and G. Csurka, “Metric learning for large scale image classification: Generalizing to new classes at near-zero cost,” in ECCV 2012-12th European Conference on Computer Vision , vol. 7573. Springer, 2012, pp. 488–501
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik, “Simultaneous detection and segmentation,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13 . Springer, 2014, pp. 297–312
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3431–3440
2015
Earlier work this paper cites.
L. Chen, J. Shen, W. Wang, and B. Ni, “Video object segmentation via dense trajectories,” IEEE Transactions on Multimedia , vol. 17, no. 12, pp. 2225–2234, 2015
2015
Earlier work this paper cites.
D. Yogatama, D. Gillick, and N. Lazic, “Embedding methods for fine grained entity type classification,” in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers) , 2015, pp. 291–296
2015
Earlier work this paper cites.
Z. Akata, F. Perronnin, Z. Harchaoui, and C. Schmid, “Label-embedding for image classification,” IEEE transactions on pattern analysis and machine intelligence , vol. 38, no. 7, pp. 1425–1438, 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, pp. 211–252, 2015
2015
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Computer Vision–ACCV 2016 Workshops: ACCV 2016 International Workshops, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13 . Springer, 2017, pp. 251–263
2017
Earlier work this paper cites.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 4, pp. 834–848, 2017
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2881–2890
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
S. Caelles, K.-K. Maninis, J. Pont-Tuset, L. Leal-Taixé, D. Cremers, and L. Van Gool, “One-shot video object segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 221–230
2017
Earlier work this paper cites.
F. Perazzi, A. Khoreva, R. Benenson, B. Schiele, and A. Sorkine-Hornung, “Learning video object segmentation from static images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2663–2672
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 609–617
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 618–626
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “Cnn architectures for large-scale audio classification,” in 2017 ieee international conference on acoustics, speech and signal processing (icassp) . IEEE, 2017, pp. 131–135
2017
Earlier work this paper cites.
S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 8759–8768
2018
Earlier work this paper cites.
S. W. Oh, J.-Y. Lee, K. Sunkavalli, and S. J. Kim, “Fast video object segmentation by reference-guided mask propagation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7376–7385
2018
Earlier work this paper cites.
Y.-T. Hu, J.-B. Huang, and A. G. Schwing, “Videomatch: Matching based video object segmentation,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 54–70
2018
Earlier work this paper cites.
Y.-T. Hu, J.-B. Huang, and A. G. Schwing, “Unsupervised video object segmentation using motion saliency-guided spatio-temporal propagation,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 786–802
2018
Earlier work this paper cites.
S. Li, B. Seybold, A. Vorobyov, A. Fathi, Q. Huang, and C.-C. J. Kuo, “Instance embedding transfer to unsupervised video object segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 6526–6535
2018
Earlier work this paper cites.
K. Gavrilyuk, A. Ghodrati, Z. Li, and C. G. Snoek, “Actor and action video segmentation from a sentence,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 5958–5966
2018
Earlier work this paper cites.
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, “Learning to localize sound source in visual scenes,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 4358–4366
2018
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Objects that sound,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 435–451
2018
Earlier work this paper cites.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” 2018
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Ulyanov, A. Vedaldi, and V. Lempitsky, “Deep image prior,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 9446–9454
2018
Earlier work this paper cites.
A. Chattopadhay, A. Sarkar, P. Howlader, and V. N. Balasubramanian, “Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks,” in 2018 IEEE winter conference on applications of computer vision (WACV) . IEEE, 2018, pp. 839–847
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7794–7803
2018
Earlier work this paper cites.
A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár, “Panoptic segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 9404–9413
2019
Earlier work this paper cites.
N. Khosravan, S. Ardeshir, and R. Puri, “On attention modules for audio-visual synchronization.” in CVPR Workshops , 2019, pp. 25–28
2019
Earlier work this paper cites.
Y. Xiong, R. Liao, H. Zhao, R. Hu, M. Bai, E. Yumer, and R. Urtasun, “Upsnet: A unified panoptic segmentation network,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8818–8826
2019
Earlier work this paper cites.
P. Voigtlaender, Y. Chai, F. Schroff, H. Adam, B. Leibe, and L.-C. Chen, “Feelvos: Fast end-to-end embedding learning for video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 9481–9490
2019
Earlier work this paper cites.
J. Johnander, M. Danelljan, E. Brissman, F. S. Khan, and M. Felsberg, “A generative appearance model for end-to-end video object segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 8953–8962
2019
Cited alongside, same era.
W. Wang, H. Song, S. Zhao, J. Shen, S. Zhao, S. C. Hoi, and H. Ling, “Learning unsupervised video object segmentation through visual attention,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 3064–3074
2019
Cited alongside, same era.
X. Lu, W. Wang, C. Ma, J. Shen, L. Shao, and F. Porikli, “See more, know more: Unsupervised video object segmentation with co-attention siamese networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 3623–3632
2019
Cited alongside, same era.
H. Wang, C. Deng, J. Yan, and D. Tao, “Asymmetric cross-guided attention network for actor and action video segmentation from natural language query,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 3939–3948
Y. Mao, J. Zhang, Z. Wan, Y. Dai, A. Li, Y. Lv, X. Tian, D.-P. Fan, and N. Barnes, “Transformer transforms salient object detection and camouflaged object detection,” 2021
2021
Later among the works it cites.
J. Zhang, J. Xie, N. Barnes, and P. Li, “Learning generative vision transformer with energy-based latent space for saliency prediction,” vol. 34, 2021, pp. 15 448–15 463
2021
Later among the works it cites.
G. Gao, G. Xu, J. Li, Y. Yu, H. Lu, and J. Yang, “Fbsnet: A fast bilateral symmetrical network for real-time semantic segmentation,” IEEE Transactions on Multimedia , 2022
2022
Later among the works it cites.
X. Yin, D. Min, Y. Huo, and S.-E. Yoon, “Contour-aware equipotential earning for semantic segmentation,” IEEE Transactions on Multimedia , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
D. Hu, F. Nie, and X. Li, “Deep multimodal clustering for unsupervised audiovisual learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 9248–9257
2019
Cited alongside, same era.
N. Pappas and J. Henderson, “Gile: A generalized input-label embedding for text classification,” Transactions of the Association for Computational Linguistics , vol. 7, pp. 139–155, 2019
2019
Cited alongside, same era.
C. Du, Z. Chen, F. Feng, L. Zhu, T. Gan, and L. Nie, “Explicit interaction model towards text classification,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 6359–6366
2019
Cited alongside, same era.
W. Huang, E. Chen, Q. Liu, Y. Chen, Z. Huang, Y. Liu, Z. Zhao, D. Zhang, and S. Wang, “Hierarchical multi-label text classification: An attention-based recurrent network approach,” in Proceedings of the 28th ACM international conference on information and knowledge management , 2019, pp. 1051–1060
2019
Cited alongside, same era.
Z. Cheng, M. Gadelha, S. Maji, and D. Sheldon, “A bayesian perspective on the deep image prior,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 5443–5451
2019
Cited alongside, same era.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
F. Z. Kaghat, A. Azough, M. Fakhour, and M. Meknassi, “A new audio augmented reality interaction and adaptation model for museum visits,” Computers & Electrical Engineering , vol. 84, p. 106606, 2020
2020
Cited alongside, same era.
C. Gan, Y. Zhang, J. Wu, B. Gong, and J. B. Tenenbaum, “Look, listen, and act: Towards audio-visual embodied navigation,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 9701–9707
2020
Cited alongside, same era.
M. Li, W. Cai, K. Verspoor, S. Pan, X. Liang, and X. Chang, “Cross-modal clinical graph transformer for ophthalmic report generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 20 656–20 665
2022
Later among the works it cites.
Z. Yang, J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. Torr, “Lavt: Language-aware vision transformer for referring image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 155–18 165
2022
Later among the works it cites.
T. Lüddecke and A. Ecker, “Image segmentation using text and image prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 7086–7096
2022
Later among the works it cites.
J. Zhou, J. Wang, J. Zhang, W. Sun, J. Zhang, S. Birchfield, D. Guo, L. Kong, M. Wang, and Y. Zhong, “Audio–visual segmentation,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVII . Springer, 2022, pp. 386–403
2022
Later among the works it cites.
S. H. Lee, W. Roh, W. Byeon, S. H. Yoon, C. Kim, J. Kim, and S. Kim, “Sound-guided semantic image manipulation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 3377–3386
2022
Later among the works it cites.
S. H. Lee, G. Oh, W. Byeon, C. Kim, W. J. Ryoo, S. H. Yoon, H. Cho, J. Bae, J. Kim, and S. Kim, “Sound-guided semantic video generation,” in European Conference on Computer Vision . Springer, 2022, pp. 34–50
2022
Later among the works it cites.
T.-J. Fu, X. E. Wang, S. T. Grafton, M. P. Eckstein, and W. Y. Wang, “M3l: Language-based video editing via multi-modal multi-level transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 513–10 522
2022
Later among the works it cites.
J. Yang, A. Barde, and M. Billinghurst, “Audio augmented reality: A systematic review of technologies, applications, and future research directions,” journal of the audio engineering society , vol. 70, no. 10, pp. 788–809, 2022
2022
Later among the works it cites.
L. Ru, Y. Zhan, B. Yu, and B. Du, “Learning affinity from attention: end-to-end weakly-supervised semantic segmentation with transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 846–16 855
2022
Later among the works it cites.
T. Zhang, S. Wei, and S. Ji, “E2ec: An end-to-end contour-based method for high-quality high-speed instance segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4443–4452
2022
Later among the works it cites.
W. Chen, D. Hong, Y. Qi, Z. Han, S. Wang, L. Qing, Q. Huang, and G. Li, “Multi-attention network for compressed video referring object segmentation,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 4416–4425
2022
Later among the works it cites.
Z. Ding, T. Hui, J. Huang, X. Wei, J. Han, and S. Liu, “Language-bridged spatial-temporal interaction for referring video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4964–4973
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Botach, E. Zheltonozhskii, and C. Baskin, “End-to-end referring video object segmentation with multimodal transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4985–4995
2022
Later among the works it cites.
J. Wu, Y. Jiang, P. Sun, Z. Yuan, and P. Luo, “Language as queries for referring video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4974–4984
2022
Later among the works it cites.
Z. Song, Y. Wang, J. Fan, T. Tan, and Z. Zhang, “Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 3222–3231
2022
Later among the works it cites.
T. Afouras, Y. M. Asano, F. Fagan, A. Vedaldi, and F. Metze, “Self-supervised object detection from audio-visual correspondence,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 575–10 586
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Chen and S. Lv, “Long text truncation algorithm based on label embedding in text classification,” Applied Sciences , vol. 12, no. 19, p. 9874, 2022
2022
Later among the works it cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pvt v2: Improved baselines with pyramid vision transformer,” Computational Visual Media , vol. 8, no. 3, pp. 415–424, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
F. Liu, Y. Liu, Y. Kong, K. Xu, L. Zhang, B. Yin, G. Hancke, and R. Lau, “Referring image segmentation using text supervision,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 22 124–22 134
2023
Closest in time.
A. Younes, D. Honerkamp, T. Welschehold, and A. Valada, “Catch me if you hear me: Audio-visual navigation in complex unmapped environments with moving sounds,” IEEE Robotics and Automation Letters , vol. 8, no. 2, pp. 928–935, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Mao, J. Zhang, M. Xiang, Y. Zhong, and Y. Dai, “Multimodal variational auto-encoder based audio-visual segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 954–965
2023
Closest in time.
C. Liu, P. P. Li, X. Qi, H. Zhang, L. Li, D. Wang, and X. Yu, “Audio-visual segmentation by exploring cross-modal mutual semantics,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 7590–7598
2023
Closest in time.
2023
Closest in time.
C. Lyu, W. Li, T. Ji, L. Wang, L. Zhou, C. Gurrin, L. Yang, Y. Yu, Y. Graham, and J. Foster, “Graph-based video-language learning with multi-grained audio-visual alignment,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 3975–3984
2023
Closest in time.
B. Miao, M. Bennamoun, Y. Gao, and A. Mian, “Spectrum-guided multi-granularity referring video object segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 920–930
2023
Closest in time.
Y.-B. Lin, H.-Y. Tseng, H.-Y. Lee, Y.-Y. Lin, and M.-H. Yang, “Unsupervised sound localization via iterative contrastive learning,” Computer Vision and Image Understanding , vol. 227, p. 103602, 2023
2023
Closest in time.
Q. Yang, X. Nie, T. Li, P. Gao, Y. Guo, C. Zhen, P. Yan, and S. Xiang, “Cooperation does matter: Exploring multi-order bilateral relations for audio-visual segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 27 134–27 143
2023
Closest in time.
M. Lan, F. Rong, Z. Li, W. Yu, and L. Zhang, “Bidirectional correlation-driven inter-frame interaction transformer for referring video object segmentation,” Pattern Recognition , vol. 153, p. 110535, 2024
2024
Closest in time.
Z. Luo, Y. Xiao, Y. Liu, S. Li, Y. Wang, Y. Tang, X. Li, and Y. Yang, “Soc: Semantic-assisted object cluster for referring video object segmentation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Chen, Y. Liu, H. Wang, F. Liu, C. Wang, H. Frazer, and G. Carneiro, “Unraveling instance associations: A closer look for audio-visual segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 497–26 507
2024
Closest in time.
S. Gao, Z. Chen, G. Chen, W. Wang, and T. Lu, “Avsegformer: Audio-visual segmentation with transformer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 11, 2024, pp. 12 155–12 163
2024
Closest in time.