Fetching the paper…
Reading the bibliography…
Given an audio-visual pair, audio-visual segmentation (AVS) aims to locate sounding sources by predicting pixel-wise maps.
S. Bird, E. Klein, and E. Loper, Natural language processing with Python: analyzing text with the natural language toolkit . ” O’Reilly Media, Inc.”, 2009
2009
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” IJCV , vol. 88, pp. 303–338, 2010
2010
Earlier work this paper cites.
J. Förster, “Local and global cross-modal influences between vision and hearing, tasting, smelling, or touching.” Journal of Experimental Psychology: General , vol. 140, no. 3, p. 364, 2011
2011
Earlier work this paper cites.
Y.-C. Chen, S.-L. Yeh, and C. Spence, “Crossmodal constraints on human perceptual awareness: auditory semantic modulation of binocular rivalry,” Frontiers in psychology , vol. 2, p. 212, 2011
2011
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in EMNLP , 2014, pp. 1532–1543. [Online]. Available: http://www.aclweb.org/anthology/D14-1162
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in MICCAI . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” ICCV , vol. 115, pp. 211–252, 2015
2015
Earlier work this paper cites.
F. Milletari, N. Navab, and S.-A. Ahmadi, “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” in 3DV . Ieee, 2016, pp. 565–571
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , June 2016
2016
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, “Learning to localize sound source in visual scenes,” in CVPR , 2018, pp. 4358–4366
2018
Earlier work this paper cites.
J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” in CVPR , 2019, pp. 3146–3154
2019
Earlier work this paper cites.
Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, and W. Liu, “Ccnet: Criss-cross attention for semantic segmentation,” in ICCV , 2019, pp. 603–612
2019
Cited alongside, same era.
T. Oya, S. Iwase, R. Natsume, T. Itazuri, S. Yamaguchi, and S. Morishima, “Do we need sound for sound source localization?” in ACCV , 2020
2020
Cited alongside, same era.
P. Li, Y. Wei, and Y. Yang, “Consistent structural relation learning for zero-shot segmentation,” NeurIPS , vol. 33, pp. 10 317–10 327, 2020
2020
Cited alongside, same era.
——, “Meta parsing networks: Towards generalized few-shot scene parsing with adaptive metric learning,” in ACMMM , 2020, pp. 64–72
2020
Cited alongside, same era.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “Vggsound: A large-scale audio-visual dataset,” in ICASSP . IEEE, 2020, pp. 721–725
2020
Cited alongside, same era.
S. Mo and P. Morgado, “Localizing visual sounds the easy way,” in ECCV . Springer, 2022, pp. 218–234
2022
Later among the works it cites.
A. Senocak, H. Ryu, J. Kim, and I. S. Kweon, “Less can be more: Sound source localization with a classification model,” in WACV , 2022, pp. 3308–3317
2022
Later among the works it cites.
J. Shi and C. Ma, “Unsupervised sounding object localization with bottom-up and top-down attention,” in WACV , 2022, pp. 1737–1746
2022
Later among the works it cites.
Z. Song, Y. Wang, J. Fan, T. Tan, and Z. Zhang, “Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes,” in CVPR , 2022, pp. 3222–3231
2022
Later among the works it cites.
A. Senocak, H. Ryu, J. Kim, and I. S. Kweon, “Learning sound localization better from semantically similar samples,” in ICASSP , 2022, pp. 4863–4867
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Tian, D. Hu, and C. Xu, “Cyclic co-learning of sounding object visual grounding and sound separation,” in CVPR , 2021, pp. 2745–2754
2021
Cited alongside, same era.
H. Chen, W. Xie, T. Afouras, A. Nagrani, A. Vedaldi, and A. Zisserman, “Localizing visual sounds the hard way,” in CVPR , 2021, pp. 16 867–16 876
2021
Cited alongside, same era.
Y. Yuan, L. Huang, J. Guo, C. Zhang, X. Chen, and J. Wang, “Ocnet: Object context for semantic segmentation,” IJCV , vol. 129, no. 8, pp. 2375–2398, 2021
2021
Cited alongside, same era.
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr et al. , “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in CVPR , 2021, pp. 6881–6890
2021
Cited alongside, same era.
B. Cheng, A. Schwing, and A. Kirillov, “Per-pixel classification is not all you need for semantic segmentation,” NeurIPS , vol. 34, pp. 17 864–17 875, 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in ICCV , 2021, pp. 568–578
2021
Cited alongside, same era.
Later among the works it cites.
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” in CVPR , 2022, pp. 1290–1299
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Closest in time.
X. Zhou, D. Zhou, D. Hu, H. Zhou, and W. Ouyang, “Exploiting visual context semantics for sound source localization,” in WACV , 2023, pp. 5199–5208
2023
Closest in time.
J. Jain, J. Li, M. T. Chiu, A. Hassani, N. Orlov, and H. Shi, “Oneformer: One transformer to rule universal image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2989–2998
2023
Closest in time.
X. Qi, C. Liu, M. Sun, L. Li, C. Fan, and X. Yu, “Diverse 3d hand gesture prediction from body dynamics by bilateral hand disentanglement,” in CVPR , 2023, pp. 4616–4626
2023
Closest in time.
2023
Closest in time.
H. Zhang, F. Li, H. Xu, S. Huang, S. Liu, L. M. Ni, and L. Zhang, “Mp-former: Mask-piloted transformer for image segmentation,” in CVPR , 2023, pp. 18 074–18 083
2023
Closest in time.