Fetching the paper…
Reading the bibliography…
Audio-Visual Segmentation (AVS) is a challenging task, which aims to segment sounding objects in video frames by exploring audio signals.
“Odor/taste integration and the perception of flavor,”
Dana M Small and John Prescott, · 2005
Earlier work this paper cites.
“The pascal visual object classes (voc) challenge,”
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman, · 2010
Earlier work this paper cites.
“Imagenet large scale visual recognition challenge,”
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al., · 2015
Earlier work this paper cites.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Earlier work this paper cites.
“V-net: Fully convolutional neural networks for volumetric medical image segmentation,”
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Look, listen and learn,”
Relja Arandjelovic and Andrew Zisserman, · 2017
Earlier work this paper cites.
“Feature pyramid networks for object detection,”
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie, · 2017
Earlier work this paper cites.
“Cnn architectures for large-scale audio classification,”
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al., · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Focal loss for dense object detection,”
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár, · 2017
Earlier work this paper cites.
“Objects that sound,”
Relja Arandjelovic and Andrew Zisserman, · 2018
Earlier work this paper cites.
“Audio-visual event localization in unconstrained videos,”
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu, · 2018
Cited alongside, same era.
“Learning to localize sound source in visual scenes,”
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon, · 2018
Cited alongside, same era.
“Dual-modality seq2seq network for audio-visual event localization,”
Yan-Bo Lin, Yu-Jhe Li, and Yu-Chiang Frank Wang, · 2019
Cited alongside, same era.
“Dual attention matching for audio-visual event localization,”
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang, · 2019
Cited alongside, same era.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2019
Cited alongside, same era.
“See the sound, hear the pixels,”
Janani Ramaswamy and Sukhendu Das, · 2020
Cited alongside, same era.
“Exploring heterogeneous clues for weakly-supervised audio-visual video parsing,”
Yu Wu and Yi Yang, · 2021
Later among the works it cites.
“Localizing visual sounds the hard way,”
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman, · 2021
Later among the works it cites.
“Sstvos: Sparse spatiotemporal transformers for video object segmentation,”
Brendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi, and Graham W Taylor, · 2021
Later among the works it cites.
“Transformer transforms salient object detection and camouflaged object detection,”
Yuxin Mao, Jing Zhang, Zhexiong Wan, Yuchao Dai, Aixuan Li, Yunqiu Lv, Xinyu Tian, Deng-Ping Fan, and Nick Barnes, · 2021
Later among the works it cites.
“Learning generative vision transformer with energy-based latent space for saliency prediction,”
Jing Zhang, Jianwen Xie, Nick Barnes, and Ping Li, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Audiovisual transformer with instance attention for audio-visual event localization,”
Yan-Bo Lin and Yu-Chiang Frank Wang, · 2020
Cited alongside, same era.
“Cross-modal relation-aware networks for audio-visual event localization,”
Haoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan, and Chuang Gan, · 2020
Cited alongside, same era.
“Unified multisensory perception: Weakly-supervised audio-visual video parsing,”
Yapeng Tian, Dingzeyu Li, and Chenliang Xu, · 2020
Cited alongside, same era.
“Multiple sound sources localization from coarse to fine,”
Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, and Weiyao Lin, · 2020
Cited alongside, same era.
“Making a case for 3d convolutions for object segmentation in videos,”
Sabarinath Mahadevan, Ali Athar, Aljoša Ošep, Sebastian Hennen, Laura Leal-Taixé, and Bastian Leibe, · 2020
Cited alongside, same era.
“Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing,”
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin, and Ming-Hsuan Yang, · 2021
Cited alongside, same era.
“Swin transformer: Hierarchical vision transformer using shifted windows,”
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo, · 2021
Later among the works it cites.
“Joint-modal label denoising for weakly-supervised audio-visual video parsing,”
Haoyue Cheng, Zhaoyang Liu, Hang Zhou, Chen Qian, Wayne Wu, and Limin Wang, · 2022
Later among the works it cites.
“Audio–visual segmentation,”
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong, · 2022
Later among the works it cites.
“Towards open vocabulary learning: A survey,”
Jianzong Wu, Xiangtai Li, Shilin Xu Haobo Yuan, Henghui Ding, Yibo Yang, Xia Li, Jiangning Zhang, Yunhai Tong, Xudong Jiang, Bernard Ghanem, et al., · 2023
Closest in time.
“Sfnet: Faster and accurate semantic segmentation via semantic flow,”
Xiangtai Li, Jiangning Zhang, Yibo Yang, Guangliang Cheng, Kuiyuan Yang, Yunhai Tong, and Dacheng Tao, · 2023
Closest in time.
“Multimodal learning with transformers: A survey,”
Peng Xu, Xiatian Zhu, and David A Clifton, · 2023
Closest in time.