Fetching the paper…
Reading the bibliography…
We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame.
Y.-B. Lin, Y.-J. Li, and Y.-C. F. Wang, “Dual-modality seq2seq network for audio-visual event localization,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 2002–2006
2006
Earlier work this paper cites.
2008
Earlier work this paper cites.
A. Faktor and M. Irani, “Video segmentation by non-local consensus voting.” in British Machine Vision Conference (BMVC) , 2014, pp. 1–8
2014
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 3431–3440
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI) , 2015, pp. 234–241
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision (IJCV) , pp. 211–252, 2015
2015
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” Advances in Neural Information Processing Systems (NeurIPS) , 2016
2016
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 2921–2929
2016
Earlier work this paper cites.
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodology for video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 724–732
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 609–617
2017
Earlier work this paper cites.
S. Caelles, K.-K. Maninis, J. Pont-Tuset, L. Leal-Taixé, D. Cremers, and L. Van Gool, “One-shot video object segmentation,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 221–230
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 618–626
2017
Earlier work this paper cites.
P. Tokmakov, K. Alahari, and C. Schmid, “Learning motion patterns in videos,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 3386–3394
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y.-T. Hu, J.-B. Huang, and A. Schwing, “Maskrnn: Instance level video object segmentation,” Advances in neural information processing systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , pp. 834–848, 2017
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “CNN architectures for large-scale audio classification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 131–135
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Objects that sound,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 435–451
2018
Earlier work this paper cites.
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-visual event localization in unconstrained videos,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 247–263
2018
Earlier work this paper cites.
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, “Learning to localize sound source in visual scenes,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 4358–4366
2018
Earlier work this paper cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 631–648
2018
Earlier work this paper cites.
Y. Zhong, Y. Dai, and H. Li, “3d geometry-aware semantic labeling of outdoor street scenes,” in 2018 24th International Conference on Pattern Recognition (ICPR) . IEEE, 2018, pp. 2343–2349
2018
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 570–586
2018
Cited alongside, same era.
R. Gao, R. Feris, and K. Grauman, “Learning to separate object sounds by watching unlabeled video,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 35–53
2018
Cited alongside, same era.
H. Song, W. Wang, S. Zhao, J. Shen, and K.-M. Lam, “Pyramid dilated deeper convlstm for video salient object detection,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 715–731
2018
Cited alongside, same era.
Y. Chen, J. Pont-Tuset, A. Montes, and L. Van Gool, “Blazingly fast video object segmentation with pixel-wise metric learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 1189–1198
2018
Cited alongside, same era.
S. Seo, J.-Y. Lee, and B. Han, “Urvos: Unified referring video object segmentation network with a large-scale benchmark,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2020, pp. 208–223
2020
Later among the works it cites.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A large-scale audio-visual dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 721–725
2020
Later among the works it cites.
S. Mahadevan, A. Athar, A. Ošep, S. Hennen, L. Leal-Taixé, and B. Leibe, “Making a case for 3D convolutions for object segmentation in videos,” in British Machine Vision Conference (BMVC) , 2020, pp. 1–15
2020
Later among the works it cites.
J. Zhou, L. Zheng, Y. Zhong, S. Hao, and M. Wang, “Positive sample propagation along the audio-visual event line,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 8436–8444
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Khoreva, A. Rohrbach, and B. Schiele, “Video object segmentation with language referring expressions,” in Proceedings of the Asian Conference on Computer Vision (ACCV) , 2018, pp. 123–141
2018
Cited alongside, same era.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 7794–7803
2018
Cited alongside, same era.
Y. Wu, L. Zhu, Y. Yan, and Y. Yang, “Dual attention matching for audio-visual event localization,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2019, pp. 6292–6300
2019
Cited alongside, same era.
D. Hu, F. Nie, and X. Li, “Deep multimodal clustering for unsupervised audiovisual learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 9248–9257
2019
Cited alongside, same era.
A. Rouditchenko, H. Zhao, C. Gan, J. McDermott, and A. Torralba, “Self-supervised audio-visual co-segmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 2357–2361
2019
Cited alongside, same era.
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba, “The sound of motions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019, pp. 1735–1744
2019
Cited alongside, same era.
R. Gao and K. Grauman, “Co-separating sounds of visual objects,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019, pp. 3879–3888
2019
Cited alongside, same era.
C. Ventura, M. Bellver, A. Girbau, A. Salvador, F. Marques, and X. Giro-i Nieto, “Rvos: End-to-end recurrent network for video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 5277–5286
2019
Cited alongside, same era.
Y. Wu and Y. Yang, “Exploring heterogeneous clues for weakly-supervised audio-visual video parsing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 1326–1335
2021
Later among the works it cites.
Y.-B. Lin, H.-Y. Tseng, H.-Y. Lee, Y.-Y. Lin, and M.-H. Yang, “Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing,” Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
H. Chen, W. Xie, T. Afouras, A. Nagrani, A. Vedaldi, and A. Zisserman, “Localizing visual sounds the hard way,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 16 867–16 876
2021
Later among the works it cites.
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” in Advances in Neural Information Processing Systems (NeurIPS) , 2021, pp. 12 077–12 090
2021
Later among the works it cites.
B. Duke, A. Ahmed, C. Wolf, P. Aarabi, and G. W. Taylor, “SSTVOS: Sparse spatiotemporal transformers for video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 5912–5921
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Zhang, J. Xie, N. Barnes, and P. Li, “Learning generative vision transformer with energy-based latent space for saliency prediction,” Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
Z. Yang, Y. Wei, and Y. Yang, “Associating objects with transformers for video object segmentation,” in Advances in Neural Information Processing Systems (NeurIPS) , 2021, pp. 1–20
2021
Later among the works it cites.
2022
Later among the works it cites.
J. Zhou, D. Guo, and M. Wang, “Contrastive positive sample propagation along the audio-visual event line,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , pp. 1–18, 2022
2022
Later among the works it cites.
H. Wang, Z.-J. Zha, L. Li, X. Chen, and J. Luo, “Semantic and relation modulation for audio-visual event localization,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , pp. 1–15, 2022
2022
Later among the works it cites.
J. Yu, Y. Cheng, R.-W. Zhao, R. Feng, and Y. Zhang, “MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video parsing,” Proceedings of the 30th ACM International Conference on Multimedia (ACM MM) , 2022
2022
Later among the works it cites.
X. Jiang, X. Xu, Z. Chen, J. Zhang, J. Song, F. Shen, H. Lu, and H. T. Shen, “Dhhn: Dual hierarchical hybrid network for weakly-supervised audio-visual video parsing,” in Proceedings of the 30th ACM International Conference on Multimedia (ACM MM) , 2022, pp. 719–727
2022
Later among the works it cites.
H. Cheng, Z. Liu, H. Zhou, C. Qian, W. Wu, and L. Wang, “Joint-modal label denoising for weakly-supervised audio-visual video parsing,” Proceedings of the European Conference on Computer Vision (ECCV) , pp. 431–448, 2022
2022
Later among the works it cites.
S. Mo and Y. Tian, “Multi-modal grouping network for weakly-supervised audio-visual video parsing,” in Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Later among the works it cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “PVTv2: Improved baselines with pyramid vision transformer,” Computational Visual Media , pp. 1–10, 2022
2022
Later among the works it cites.
J. Zhou, J. Wang, J. Zhang, W. Sun, J. Zhang, S. Birchfield, D. Guo, L. Kong, M. Wang, and Y. Zhong, “Audio–visual segmentation,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2022, pp. 386–403
2022
Later among the works it cites.
J. Wu, Y. Jiang, P. Sun, Z. Yuan, and P. Luo, “Language as queries for referring video object segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 4974–4984
2022
Later among the works it cites.
A. Botach, E. Zheltonozhskii, and C. Baskin, “End-to-end referring video object segmentation with multimodal transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 4985–4995
2022
Later among the works it cites.