Fetching the paper…
Reading the bibliography…
Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video.
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia. Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Weakly-and semi-supervised learning of a deep convolutional network for semantic image segmentation
George Papandreou, Liang-Chieh Chen, Kevin P Murphy, and Alan L Yuille · 2015
Earlier work this paper cites.
Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation
Jifeng Dai, Kaiming He, and Jian Sun · 2015
Earlier work this paper cites.
Scribblesup: Scribble-supervised convolutional networks for semantic segmentation
Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun · 2016
Earlier work this paper cites.
What’s the point: Semantic segmentation with point supervision
Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2016
Earlier work this paper cites.
Seed, expand and constrain: Three principles for weakly-supervised image segmentation
Alexander Kolesnikov and Christoph H Lampert · 2016
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H. McDermott, William T. Freeman, and Antonio Torralba · 2016
Earlier work this paper cites.
Audio event and scene recognition: A unified approach using strongly and weakly labeled data
Anurag Kumar and Bhiksha Raj · 2016
Earlier work this paper cites.
Audio event detection using weakly labeled data
Anurag Kumar and Bhiksha Raj · 2016
Earlier work this paper cites.
Deep clustering: Discriminative embeddings for segmentation and separation
John R. Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning random-walk label propagation for weakly-supervised semantic segmentation
Paul Vernaza and Manmohan Chandraker · 2017
Earlier work this paper cites.
Bottom-up top-down cues for weakly-supervised semantic segmentation
Qibin Hou, Daniela Massiceti, Puneet Kumar Dokania, Yunchao Wei, Ming-Ming Cheng, and Philip HS Torr · 2017
Earlier work this paper cites.
Object region mining with adversarial erasing: A simple classification to semantic segmentation approach
Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan · 2017
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Simple does it: Weakly supervised instance and semantic segmentation
Anna Khoreva, Rodrigo Benenson, Jan Hosang, Matthias Hein, and Bernt Schiele · 2017
Earlier work this paper cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Earlier work this paper cites.
Weakly-supervised semantic segmentation network with deep seeded region growing
Zilong Huang, Xinggang Wang, Jiasi Wang, Wenyu Liu, and Jingdong Wang · 2018
Earlier work this paper cites.
Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation
Yunchao Wei, Huaxin Xiao, Honghui Shi, Zequn Jie, Jiashi Feng, and Thomas S Huang · 2018
Earlier work this paper cites.
Tell me where to look: Guided attention inference network
Kunpeng Li, Ziyan Wu, Kuan-Chuan Peng, Jan Ernst, and Yun Fu · 2018
Earlier work this paper cites.
Self-erasing network for integral object attention
Qibin Hou, PengTao Jiang, Yunchao Wei, and Ming-Ming Cheng · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Earlier work this paper cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Earlier work this paper cites.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Earlier work this paper cites.
Self-supervised generation of spatial audio for 360°video
Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang · 2018
Cited alongside, same era.
A closer look at weak label learning for audio events
Ankit Shah, Anurag Kumar, Alexander Hauptmann, and Bhiksha Raj · 2018
Cited alongside, same era.
Self-supervised audio-visual co-segmentation
Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh H. McDermott, and Antonio Torralba · 2019
Cited alongside, same era.
Deep multimodal clustering for unsupervised audiovisual learning
Di Hu, Feiping Nie, and Xuelong Li · 2019
Cited alongside, same era.
Ficklenet: Weakly and semi-supervised semantic image segmentation using stochastic inference
Jungbeom Lee, Eunji Kim, Sungmin Lee, Jangho Lee, and Sungroh Yoon · 2019
Cited alongside, same era.
Differentiable multi-granularity human representation learning for instance-aware human semantic parsing
Tianfei Zhou, Wenguan Wang, Si Liu, Yi Yang, and Luc Van Gool · 2021
Later among the works it cites.
Deep graph cut network for weakly-supervised semantic segmentation
Jiapei Feng, Xinggang Wang, and Wenyu Liu · 2021
Later among the works it cites.
Robust audio-visual instance discrimination
Pedro Morgado, Ishan Misra, and Nuno Vasconcelos · 2021
Later among the works it cites.
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2021
Later among the works it cites.
Cyclic co-learning of sounding object visual grounding and sound separation
Yapeng Tian, Di Hu, and Chenliang Xu · 2021
Later among the works it cites.
Visualvoice: Audio-visual speech separation with cross-modal consistency
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wataru Shimoda and Keiji Yanai · 2019
Cited alongside, same era.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Cited alongside, same era.
Co-separating sounds of visual objects
Ruohan Gao and Kristen Grauman · 2019
Cited alongside, same era.
2.5d visual sound
Ruohan Gao and Kristen Grauman · 2019
Cited alongside, same era.
Class-conditional embeddings for music source separation
Prem Seetharaman, Gordon Wichern, Shrikant Venkataramani, and Jonathan Le Roux · 2019
Cited alongside, same era.
Finding strength in weakness: Learning to separate sounds with weak supervision
Fatemeh Pishdadian, Gordon Wichern, and Jonathan Le Roux · 2019
Cited alongside, same era.
Panoptic feature pyramid networks
Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollar · 2019
Cited alongside, same era.
Ruohan Gao and Kristen Grauman · 2021
Later among the works it cites.
Weakly-supervised audio-visual sound source detection and separation
Tanzila Rahman and Leonid Sigal · 2021
Later among the works it cites.
Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Yu Wu and Yi Yang · 2021
Later among the works it cites.
Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin, and Ming-Hsuan Yang · 2021
Later among the works it cites.
Multiple instance graph learning for weakly supervised remote sensing object detection
Binglu Wang, Yongqiang Zhao, and Xuelong Li · 2021
Later among the works it cites.
Background-aware pooling and noise-aware loss for weakly-supervised semantic segmentation
Youngmin Oh, Beomjun Kim, and Bumsub Ham · 2021
Later among the works it cites.
Audio-visual segmentation
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong · 2022
Later among the works it cites.
Localizing visual sounds the easy way
Shentong Mo and Pedro Morgado · 2022
Later among the works it cites.
A closer look at weakly-supervised audio-visual source localization
Shentong Mo and Pedro Morgado · 2022
Later among the works it cites.
C2am: Contrastive learning of class-agnostic activation map for weakly supervised object localization and semantic segmentation
Jinheng Xie, Jianfeng Xiang, Junliang Chen, Xianxu Hou, Xiaodong Zhao, and Linlin Shen · 2022
Later among the works it cites.
Semantic-aware multi-modal grouping for weakly-supervised audio-visual video parsing
Shentong Mo and Yapeng Tian · 2022
Later among the works it cites.
Benchmarking weakly-supervised audio-visual sound localization
Shentong Mo and Pedro Morgado · 2022
Later among the works it cites.
Multi-modal grouping network for weakly-supervised audio-visual video parsing
Shentong Mo and Yapeng Tian · 2022
Later among the works it cites.
Learning sound localization better from semantically similar samples
Arda Senocak, Hyeonggon Ryu, Junsik Kim, and In So Kweon · 2022
Later among the works it cites.
DiffAVA: Personalized text-to-audio generation with visual alignment
Shentong Mo, Jing Shi, and Yapeng Tian · 2023
Closest in time.
A unified audio-visual learning framework for localization, separation, and recognition
Shentong Mo and Pedro Morgado · 2023
Closest in time.
Audio-visual class-incremental learning
Weiguo Pian, Shentong Mo, Yunhui Guo, and Yapeng Tian · 2023
Closest in time.
Class-incremental grouping network for continual audio-visual learning
Shentong Mo, Weiguo Pian, and Yapeng Tian · 2023
Closest in time.
Audio-visual grouping network for sound localization from mixtures
Shentong Mo and Yapeng Tian · 2023
Closest in time.
AV-SAM: Segment anything model meets audio-visual localization and segmentation
Shentong Mo and Yapeng Tian · 2023
Closest in time.