Fetching the paper…
Reading the bibliography…
Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks.
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia. Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Self-supervised generation of spatial audio for 360°video
Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Learning representations from audio-visual spatial alignment
Pedro Morgado, Yi Li, and Nuno Nvasconcelos · 2020
Cited alongside, same era.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu · 2020
Cited alongside, same era.
A closer look at weakly-supervised audio-visual source localization
Shentong Mo and Pedro Morgado · 2022
Cited alongside, same era.
Localizing visual sounds the easy way
Shentong Mo and Pedro Morgado · 2022
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun
Cited in the paper.
Multi-modal grouping network for weakly-supervised audio-visual video parsing
Shentong Mo and Yapeng Tian · 2022
Later among the works it cites.
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong · 2022
Later among the works it cites.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Closest in time.
Audio-visual grouping network for sound localization from mixtures
Shentong Mo and Yapeng Tian · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…