Fetching the paper…
Reading the bibliography…
Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance.
S. Roweis, “One microphone source separation,” Advances in neural information processing systems , vol. 13, 2000
2000
Earlier work this paper cites.
J. Hershey and M. Casey, “Audio-visual sound separation via hidden markov models,” Advances in Neural Information Processing Systems , vol. 14, 2001
2001
Earlier work this paper cites.
P. Smaragdis and J. Brown, “Non-negative matrix factorization for polyphonic music transcription,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , 2003, pp. 177–180
2003
Earlier work this paper cites.
Y.-B. Lin, Y.-J. Li, and Y.-C. F. Wang, “Dual-modality seq2seq network for audio-visual event localization,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 2002–2006
2006
Earlier work this paper cites.
T. Virtanen, “Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 3, pp. 1066–1074, 2007
2007
Earlier work this paper cites.
L. van der Maaten and G. Hinton, “Visualizing data using t-sne,” Journal of Machine Learning Research , vol. 9, no. 86, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
A. Cichocki, R. Zdunek, A. H. Phan, and S.-i. Amari, Nonnegative matrix and tensor factorizations: applications to exploratory multi-way data analysis and blind source separation . John Wiley & Sons, 2009
2009
Earlier work this paper cites.
G. J. Mysore, P. Smaragdis, and B. Raj, “Non-negative hidden markov modeling of audio with application to source separation,” in International Conference on Latent Variable Analysis and Signal Separation . Springer, 2010, pp. 140–148
2010
Earlier work this paper cites.
P.-S. Huang, S. D. Chen, P. Smaragdis, and M. Hasegawa-Johnson, “Singing-voice separation from monaural recordings using robust principal component analysis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2012, pp. 57–60
2012
Earlier work this paper cites.
C. Raffel, B. Mcfee, E. Humphrey, J. Salamon, O. Nieto, D. Liang, and D. Ellis, “mir_eval: A transparent implementation of common mir metrics,” in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR) , 2014
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2016
2016
Earlier work this paper cites.
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba, “Ambient sound provides supervision for visual learning,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2016, pp. 801–816
2016
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 31–35
2016
Earlier work this paper cites.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “Yfcc100m: The new data in multimedia research,” Commun. ACM , vol. 59, no. 2, p. 64–73, 2016
2016
Earlier work this paper cites.
A. Kumar and B. Raj, “Audio event and scene recognition: A unified approach using strongly and weakly labeled data,” 2017 International Joint Conference on Neural Networks (IJCNN) , pp. 3475–3482, 2016
2016
Earlier work this paper cites.
——, “Audio event detection using weakly labeled data,” in Proceedings of the 24th ACM International Conference on Multimedia , 2016, p. 1038–1047
2016
Earlier work this paper cites.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “The MUSDB18 corpus for music separation,” Dec. 2017. [Online]. Available: https://doi.org/10.5281/zenodo.1117372
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 609–617
2017
Earlier work this paper cites.
E. Fonseca, J. Pons, X. Favory, F. Font, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra, “Freesound datasets: a platform for the creation of open audio datasets,” in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR) , 2017, pp. 486–493
2017
Earlier work this paper cites.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” in Proceedings of International Conference on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Ulyanov, A. Vedaldi, and V. Lempitsky, “Deep image prior,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Earlier work this paper cites.
D. Stoller, S. Ewert, and S. Dixon, “Wave-u-net: A multi-scale neural network for end-to-end audio source separation,” in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR) , 2018, pp. 334–340
2018
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 570–586
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
R. Gao, R. Feris, and K. Grauman, “Learning to separate object sounds by watching unlabeled video,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 35–53
2018
Cited alongside, same era.
B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,” in Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2018
2018
Cited alongside, same era.
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, “Learning to localize sound source in visual scenes,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 4358–4366
2018
Cited alongside, same era.
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-visual event localization in unconstrained videos,” in Proceedings of European Conference on Computer Vision (ECCV) , 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
Q. Kong, Y. Cao, H. Liu, K. Choi, and Y. Wang, “Decoupling magnitude and phase estimation with deep resunet for music source separation,” in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR) , 2021
2021
Later among the works it cites.
S. Wisdom, H. Erdogan, D. P. W. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. R. Hershey, “What’s all the fuss about free universal sound separation data?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 186–190
2021
Later among the works it cites.
Y. Tian, D. Hu, and C. Xu, “Cyclic co-learning of sounding object visual grounding and sound separation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 2745–2754
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Morgado, N. Nvasconcelos, T. Langlois, and O. Wang, “Self-supervised generation of spatial audio for 360°video,” in Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
X. Xu, B. Dai, and D. Lin, “Recursive visual sound separation using minus-plus net,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019
2019
Cited alongside, same era.
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba, “The sound of motions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019, pp. 1735–1744
2019
Cited alongside, same era.
Y. Wu, L. Zhu, Y. Yan, and Y. Yang, “Dual attention matching for audio-visual event localization,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2019, pp. 6291–6299
2019
Cited alongside, same era.
R. Gao and K. Grauman, “2.5d visual sound,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 324–333
2019
Cited alongside, same era.
D. Hu, F. Nie, and X. Li, “Deep multimodal clustering for unsupervised audiovisual learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 9248–9257
2019
Cited alongside, same era.
2021
Later among the works it cites.
P. Morgado, I. Misra, and N. Vasconcelos, “Robust audio-visual instance discrimination,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 12 934–12 945
2021
Later among the works it cites.
P. Morgado, N. Vasconcelos, and I. Misra, “Audio-visual instance discrimination with cross-modal agreement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2021, pp. 12 475–12 486
2021
Later among the works it cites.
Y. Wu and Y. Yang, “Exploring heterogeneous clues for weakly-supervised audio-visual video parsing,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 1326–1335
2021
Later among the works it cites.
Y.-B. Lin, H.-Y. Tseng, H.-Y. Lee, Y.-Y. Lin, and M.-H. Yang, “Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing,” in Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
H. Chen, W. Xie, T. Afouras, A. Nagrani, A. Vedaldi, and A. Zisserman, “Localizing visual sounds the hard way,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 16 867–16 876
2021
Later among the works it cites.
M. Chatterjee, J. Le Roux, N. Ahuja, and A. Cherian, “Visual scene graphs for audio source separation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 1204–1213
2021
Later among the works it cites.
S. Mo and Y. Tian, “Multi-modal grouping network for weakly-supervised audio-visual video parsing,” in Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Later among the works it cites.
——, “Semantic-aware multi-modal grouping for weakly-supervised audio-visual video parsing,” in European Conference on Computer Vision (ECCV) Workshop , 2022
2022
Later among the works it cites.
S. Mo and P. Morgado, “Localizing visual sounds the easy way,” in Proceedings of European Conference on Computer Vision (ECCV) , 2022, p. 218–234
2022
Later among the works it cites.
——, “A closer look at weakly-supervised audio-visual source localization,” in Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Later among the works it cites.
——, “Benchmarking weakly-supervised audio-visual sound localization,” in European Conference on Computer Vision (ECCV) Workshop , 2022
2022
Later among the works it cites.
K. Patterson, K. Wilson, S. Wisdom, and J. R. Hershey, “Distance-based sound separation,” in Proceedings of Interspeech , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Mo, W. Pian, and Y. Tian, “Class-incremental grouping network for continual audio-visual learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Later among the works it cites.
W. Pian, S. Mo, Y. Guo, and Y. Tian, “Audio-visual class-incremental learning,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Mo and B. Raj, “Weakly-supervised audio-visual segmentation,” in Proceedings of Advances in Neural Information Processing Systems (NeurIPS) , 2023
2023
Later among the works it cites.