Fetching the paper…
Reading the bibliography…
The Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Objects that sound
Relja Arandjelovic and Andrew Zisserman · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Dual-modality seq2seq network for audio-visual event localization
Yan-Bo Lin, Yu-Jhe Li, and Yu-Chiang Frank Wang · 2019
Earlier work this paper cites.
Dual attention matching for audio-visual event localization
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang · 2019
Earlier work this paper cites.
VGGSound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classification and retrieval of videos
Kranti Parida, Neeraj Matiyali, Tanaya Guha, and Gaurav Sharma · 2020
Earlier work this paper cites.
Multiple sound sources localization from coarse to fine
Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, and Weiyao Lin · 2020
Earlier work this paper cites.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu · 2020
Earlier work this paper cites.
Cross-modal relation-aware networks for audio-visual event localization
Haoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan, and Chuang Gan · 2020
Earlier work this paper cites.
Localizing visual sounds the hard way
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman · 2021
Earlier work this paper cites.
Audio-visual event localization via recursive fusion by joint co-attention
Bin Duan, Hao Tang, Wei Wang, Ziliang Zong, Guowei Yang, and Yan Yan · 2021
Earlier work this paper cites.
Avgzslnet: Audio-visual generalized zero-shot learning by reconstructing label features from multi-modal embeddings
Pratik Mazumder, Pravendra Singh, Kranti Kumar Parida, and Vinay P Namboodiri · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Positive sample propagation along the audio-visual event line
Jinxing Zhou, Liang Zheng, Yiran Zhong, Shijie Hao, and Meng Wang · 2021
Cited alongside, same era.
Learning to answer questions in dynamic audio-visual scenarios
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu · 2022
Cited alongside, same era.
Localizing visual sounds the easy way
Shentong Mo and Pedro Morgado · 2022
Cited alongside, same era.
Span-based audio-visual localization
Yiling Wu, Xinfeng Zhang, Yaowei Wang, and Qingming Huang · 2022
Cited alongside, same era.
Cross-modal background suppression for audio-visual event localization
Yan Xia and Zhou Zhao · 2022
Ave-clip: Audioclip-based multi-window temporal transformer for audio visual event localization
Tanvir Mahmud and Diana Marculescu · 2023
Later among the works it cites.
Pg-video-llava: Pixel grounding large video-language models
Shehan Munasinghe, Rusiru Thushara, Muhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Mubarak Shah, and Fahad Khan · 2023
Later among the works it cites.
Fine-grained audible video description
Xuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin, Bowen He, Xiaodong Han, Aixuan Li, Yuchao Dai, Lingpeng Kong, Meng Wang, et al · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Later among the works it cites.
Cm-pie: Cross-modal perception for interactive-enhanced audio-visual video parsing
Yaru Chen, Ruohao Guo, Xubo Liu, Peipei Wu, Guangyao Li, Zhenbo Li, and Wenwu Wang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Avqa: A dataset for audio-visual question answering on videos
Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu · 2022
Cited alongside, same era.
MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video parsing
Jiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng, and Yuejie Zhang · 2022
Cited alongside, same era.
Cross-modal label contrastive learning for unsupervised audio-visual event localization
Peijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, and Alex C Kot · 2023
Cited alongside, same era.
Collecting cross-modal presence-absence evidence for weakly-supervised audio-visual event perception
Junyu Gao, Mengyuan Chen, and Changsheng Xu · 2023
Cited alongside, same era.
Learning event-specific localization preferences for audio-visual event localization
Shiping Ge, Zhiwei Jiang, Yafeng Yin, Cong Wang, Zifeng Cheng, and Qing Gu · 2023
Cited alongside, same era.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Cited alongside, same era.
Closest in time.
Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al · 2024
Closest in time.
Improving audio-visual segmentation with bidirectional generation
Dawei Hao, Yuxin Mao, Bowen He, Xiaodong Han, Yuchao Dai, and Yiran Zhong · 2024
Closest in time.
Xiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao, Guobin Shen, Qingqun Kong, Xin Yang, and Yi Zeng · 2024
Closest in time.
T-vsl: Text-guided visual sound source localization in mixtures
Tanvir Mahmud, Yapeng Tian, and Diana Marculescu · 2024
Closest in time.
Tavgbench: Benchmarking text to audible-video generation
Yuxin Mao, Xuyang Shen, Jing Zhang, Zhen Qin, Jinxing Zhou, Mochu Xiang, Yiran Zhong, and Yuchao Dai · 2024
Closest in time.
Audio-visual generalized zero-shot learning the easy way
Shentong Mo and Pedro Morgado · 2024
Closest in time.
OpenAVE: Moving towards open set audio-visual event localization
Jiale Yu, Baopeng Zhang, Zhu Teng, and Jianping Fan · 2024
Closest in time.
Multimodal class-aware semantic enhancement network for audio-visual video parsing
Pengcheng Zhao, Jinxing Zhou, Dan Guo, Yang Zhao, and Yanxiang Chen · 2024
Closest in time.