Fetching the paper…
Reading the bibliography…
The aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level.
Making a case for 3d convolutions for object segmentation in videos
Mahadevan, S.; Athar, A.; Ošep, A.; Hennen, S.; Leal-Taixé, L.; and Leibe, B. 2020 · 2008
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Look, listen and learn
Arandjelovic, R.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; Moore, R. C.; Plakal, M.; Platt, D.; Saurous, R. A.; Seybold, B.; et al. 2017 · 2017
Earlier work this paper cites.
Learning motion patterns in videos
Tokmakov, P.; Alahari, K.; and Schmid, C. 2017 · 2017
Earlier work this paper cites.
Objects that sound
Arandjelovic, R.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
Unsupervised deep epipolar flow for stationary or dynamic scenes
Zhong, Y.; Ji, P.; Wang, J.; Dai, Y.; and Li, H. 2019 · 2019
Earlier work this paper cites.
Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning
Cheng, Y.; Wang, R.; Pan, Z.; Feng, R.; and Zhang, Y. 2020 · 2020
Earlier work this paper cites.
Discriminative sounding objects localization via self-supervised audiovisual matching
Hu, D.; Qian, R.; Jiang, M.; Tan, X.; Wen, S.; Ding, E.; Lin, W.; and Dou, D. 2020 · 2020
Cited alongside, same era.
Multiple sound sources localization from coarse to fine
Qian, R.; Hu, D.; Dinkel, H.; Wu, M.; Xu, N.; and Lin, W. 2020 · 2020
Cited alongside, same era.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Tian, Y.; Li, D.; and Xu, C. 2020 · 2020
Cited alongside, same era.
Displacement-invariant matching cost learning for accurate optical flow estimation
Wang, J.; Zhong, Y.; Dai, Y.; Zhang, K.; Ji, P.; and Li, H. 2020 · 2020
Cited alongside, same era.
Cross-modal attention network for temporal inconsistent audio-visual event localization
Xuan, H.; Zhang, Z.; Chen, S.; Yang, J.; and Yan, Y. 2020 · 2020
Cited alongside, same era.
Localizing visual sounds the hard way
Self-supervised video object segmentation by motion grouping
Yang, C.; Lamdouar, H.; Lu, E.; Zisserman, A.; and Xie, W. 2021 · 2021
Later among the works it cites.
Learning generative vision transformer with energy-based latent space for saliency prediction
Zhang, J.; Xie, J.; Barnes, N.; and Li, P. 2021 · 2021
Later among the works it cites.
Positive sample propagation along the audio-visual event line
Zhou, J.; Zheng, L.; Zhong, Y.; Hao, S.; and Wang, M. 2021 · 2021
Later among the works it cites.
Mix and localize: Localizing sound sources in mixtures
Hu, X.; Chen, Z.; and Owens, A. 2022 · 2022
Later among the works it cites.
Pvt v2: Improved baselines with pyramid vision transformer
Wang, W.; Xie, E.; Li, X.; Fan, D.-P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; and Shao, L. 2022 · 2022
Later among the works it cites.
Displacement-Invariant Cost Computation for Stereo Matching
Zhong, Y.; Loop, C.; Byeon, W.; Birchfield, S.; Dai, Y.; Zhang, K.; Kamenev, A.; Breuel, T.; Li, H.; and Kautz, J. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, H.; Xie, W.; Afouras, T.; Nagrani, A.; Vedaldi, A.; and Zisserman, A. 2021 · 2021
Cited alongside, same era.
Sstvos: Sparse spatiotemporal transformers for video object segmentation
Duke, B.; Ahmed, A.; Wolf, C.; Aarabi, P.; and Taylor, G. W. 2021 · 2021
Cited alongside, same era.
Transformer transforms salient object detection and camouflaged object detection
Mao, Y.; Zhang, J.; Wan, Z.; Dai, Y.; Li, A.; Lv, Y.; Tian, X.; Fan, D.-P.; and Barnes, N. 2021 · 2021
Cited alongside, same era.
Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Wu, Y.; and Yang, Y. 2021 · 2021
Cited alongside, same era.
Contrastive conditional latent diffusion for audio-visual segmentation
Mao, Y.; Zhang, J.; Xiang, M.; Lv, Y.; Zhong, Y.; and Dai, Y. 2023a
Cited in the paper.
Multimodal variational auto-encoder based audio-visual segmentation
Mao, Y.; Zhang, J.; Xiang, M.; Zhong, Y.; and Dai, Y. 2023b
Cited in the paper.
Later among the works it cites.
Contrastive positive sample propagation along the audio-visual event line
Zhou, J.; Guo, D.; and Wang, M. 2022 · 2022
Later among the works it cites.
Audio–Visual Segmentation
Zhou, J.; Wang, J.; Zhang, J.; Sun, W.; Zhang, J.; Birchfield, S.; Guo, D.; Kong, L.; Wang, M.; and Zhong, Y. 2022 · 2022
Later among the works it cites.
Improving audio-visual video parsing with pseudo visual labels
Zhou, J.; Guo, D.; Zhong, Y.; and Wang, M. 2023 · 2023
Closest in time.