Fetching the paper…
Reading the bibliography…
The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both.
Visualizing data using t-SNE
Van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; Moore, R. C.; Plakal, M.; Platt, D.; Saurous, R. A.; Seybold, B.; et al. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Tian, Y.; Shi, J.; Li, B.; Duan, Z.; and Xu, C. 2018 · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Tran, D.; Wang, H.; Torresani, L.; Ray, J.; LeCun, Y.; and Paluri, M. 2018 · 2018
Earlier work this paper cites.
The sound of pixels
Zhao, H.; Gan, C.; Rouditchenko, A.; Vondrick, C.; McDermott, J.; and Torralba, A. 2018 · 2018
Earlier work this paper cites.
Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning
Cheng, Y.; Wang, R.; Pan, Z.; Feng, R.; and Zhang, Y. 2020 · 2020
Earlier work this paper cites.
Multiple sound sources localization from coarse to fine
Qian, R.; Hu, D.; Dinkel, H.; Wu, M.; Xu, N.; and Lin, W. 2020 · 2020
Earlier work this paper cites.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Tian, Y.; Li, D.; and Xu, C. 2020 · 2020
Earlier work this paper cites.
Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video Parsing
Lin, Y.-B.; Tseng, H.-Y.; Lee, H.-Y.; Lin, Y.-Y.; and Yang, M.-H. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Wu, Y.; and Yang, Y. 2021 · 2021
Cited alongside, same era.
Positive sample propagation along the audio-visual event line
Zhou, J.; Zheng, L.; Zhong, Y.; Hao, S.; and Wang, M. 2021 · 2021
Cited alongside, same era.
Joint-Modal Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
Cheng, H.; Liu, Z.; Zhou, H.; Qian, C.; Wu, W.; and Wang, L. 2022 · 2022
Cited alongside, same era.
DHHN: Dual Hierarchical Hybrid Network for Weakly-Supervised Audio-Visual Video Parsing
Jiang, X.; Xu, X.; Chen, Z.; Zhang, J.; Song, J.; Shen, F.; Lu, H.; and Shen, H. T. 2022 · 2022
Cited alongside, same era.
Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event Parser
Lai, Y.-H.; Chen, Y.-C.; and Yu-Chiang, F. W. 2023 · 2023
Later among the works it cites.
Event-specific audio-visual fusion layers: A simple and new perspective on video understanding
Senocak, A.; Kim, J.; Oh, T.-H.; Li, D.; and Kweon, I. S. 2023 · 2023
Later among the works it cites.
Fine-grained audible video description
Shen, X.; Li, D.; Zhou, J.; Qin, Z.; He, B.; Han, X.; Li, A.; Dai, Y.; Kong, L.; Wang, M.; et al. 2023 · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023 · 2023
Later among the works it cites.
Multi-Modal and Multi-Scale Temporal Fusion Architecture Search for Audio-Visual Video Parsing
Zhang, J.; and Li, W. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-modal Grouping Network for Weakly-Supervised Audio-Visual Video Parsing
Mo, S.; and Tian, Y. 2022 · 2022
Cited alongside, same era.
MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video Parsing
Yu, J.; Cheng, Y.; Zhao, R.-W.; Feng, R.; and Zhang, Y. 2022 · 2022
Cited alongside, same era.
Contrastive Positive Sample Propagation along the Audio-Visual Event Line
Zhou, J.; Guo, D.; and Wang, M. 2022 · 2022
Cited alongside, same era.
Audio–visual segmentation
Zhou, J.; Wang, J.; Zhang, J.; Sun, W.; Zhang, J.; Birchfield, S.; Guo, D.; Kong, L.; Wang, M.; and Zhong, Y. 2022 · 2022
Cited alongside, same era.
Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language Perspective
Fan, Y.; Wu, Y.; Lin, Y.; and Du, B. 2023 · 2023
Cited alongside, same era.
Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio-Visual Event Perception
Gao, J.; Chen, M.; and Xu, C. 2023 · 2023
Cited alongside, same era.
Audio-Visual Instance Segmentation
Guo, R.; Ying, X.; Chen, Y.; Niu, D.; Li, G.; Qu, L.; Qi, Y.; Zhou, J.; Xing, B.; Yue, W.; Shi, J.; Wang, Q.; Zhang, P.; and Liang, B. 2023 · 2023
Cited alongside, same era.
Zhou, J.; Guo, D.; Zhong, Y.; and Wang, M. 2023 · 2023
Later among the works it cites.
UniAV: Unified Audio-Visual Perception for Multi-Task Video Localization
Geng, T.; Wang, T.; Zhang, Y.; Duan, J.; Guan, W.; and Zheng, F. 2024 · 2024
Closest in time.
Toward Long Form Audio-Visual Video Understanding
Hou, W.; Li, G.; Tian, Y.; and Hu, D. 2024 · 2024
Closest in time.
Object-Aware Adaptive-Positivity Learning for Audio-Visual Question Answering
Li, Z.; Guo, D.; Zhou, J.; Zhang, J.; and Wang, M. 2024 · 2024
Closest in time.
TAVGBench: Benchmarking text to audible-video generation
Mao, Y.; Shen, X.; Zhang, J.; Qin, Z.; Zhou, J.; Xiang, M.; Zhong, Y.; and Dai, Y. 2024 · 2024
Closest in time.
Label-anticipated event disentanglement for audio-visual video parsing
Zhou, J.; Guo, D.; Mao, Y.; Zhong, Y.; Chang, X.; and Wang, M. 2025 · 2025
Closest in time.