Fetching the paper…
Reading the bibliography…
In music source separation, the number of sources may vary for each piece and some of the sources may belong to the same family of instruments, thus sharing timbral characteristics and making the sources more correlated.
J. R. Hershey and J. R. Movellan, “Audio vision: Using audio-visual synchrony to locate sounds,” in Advances in neural information processing systems , 2000, pp. 813–819
2000
Earlier work this paper cites.
K. Lee, “An analysis and comparison of the clarinet and viola versions of the two sonatas for clarinet (or viola) and piano Op. 120 by Johannes Brahms,” Ph.D. dissertation, University of Cincinnati, 2004
2004
Earlier work this paper cites.
E. Kidron, Y. Y. Schechner, and M. Elad, “Pixels that sound,” in 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05) , vol. 1. IEEE, 2005, pp. 88–95
2005
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Fevotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1462–1469, July 2006
2006
Earlier work this paper cites.
T. Virtanen, “Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria,” IEEE transactions on audio, speech, and language processing , vol. 15, no. 3, pp. 1066–1074, 2007
2007
Earlier work this paper cites.
Z. Barzelay and Y. Y. Schechner, “Harmony in motion,” in 2007 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2007, pp. 1–8
2007
Earlier work this paper cites.
A. Ozerov and C. Févotte, “Multichannel nonnegative matrix factorization in convolutive mixtures for audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 3, pp. 550–563, 2009
2009
Earlier work this paper cites.
J. J. Carabias-Orti, T. Virtanen, P. Vera-Candeas, N. Ruiz-Reyes, and F. J. Canadas-Quesada, “Musical instrument sound multi-excitation model for non-negative spectrogram factorization,” IEEE Journal of Selected Topics in Signal Processing , vol. 5, no. 6, pp. 1144–1158, 2011
2011
Earlier work this paper cites.
J. J. Carabias-Orti, M. Cobos, P. Vera-Candeas, and F. J. Rodríguez-Serrano, “Nonnegative signal factorization with learnt instrument models for sound source separation in close-microphone recordings,” EURASIP Journal on Advances in Signal Processing , vol. 2013, no. 1, p. 184, 2013
2013
Earlier work this paper cites.
M. Adeli, J. Rouat, and S. Molotchnikoff, “Audiovisual correspondence between musical timbre and visual shapes,” Frontiers in human neuroscience , vol. 8, p. 352, 2014
2014
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
M. Miron, J. J. Carabias-Orti, J. J. Bosch, E. Gómez, and J. Janer, “Score-informed source separation for multichannel orchestral recordings,” Journal of Electrical and Computer Engineering , vol. 2016, 2016
2016
Earlier work this paper cites.
E. M. Grais, G. Roma, A. J. Simpson, and M. Plumbley, “Combining mask estimates for single channel audio source separation using deep neural networks,” Interspeech2016 Proceedings , 2016
2016
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “SoundNet: Learning sound representations from unlabeled video,” in Advances in neural information processing systems , 2016, pp. 892–900
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
M. Miron, J. Janer Mestres, and E. Gómez Gutiérrez, “Generating data to train convolutional neural networks for classical music source separation,” in Lokki T, Pätynen J, Välimäki V, editors. Proceedings of the 14th Sound and Music Computing Conference; 2017 Jul 5-8; Espoo, Finland. Aalto: Aalto University; 2017. p. 227-33. Aalto University, 2017
2016
Earlier work this paper cites.
B. Li, K. Dinesh, Z. Duan, and G. Sharma, “See and listen: Score-informed association of sound tracks to players in chamber music performance videos,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 2906–2910
2017
Earlier work this paper cites.
D. Ramachandram and G. W. Taylor, “Deep multimodal learning: A survey on recent advances and trends,” IEEE Signal Processing Magazine , vol. 34, no. 6, pp. 96–108, 2017
2017
Earlier work this paper cites.
S. Parekh, S. Essid, A. Ozerov, N. Q. Duong, P. Pérez, and G. Richard, “Guiding audio source separation by video object information,” in Applications of Signal Processing to Audio and Acoustics (WASPAA), 2017 IEEE Workshop on . IEEE, 2017, pp. 61–65
2017
Earlier work this paper cites.
P. Chandna, M. Miron, J. Janer, and E. Gómez, “Monoaural audio source separation using deep convolutional neural networks,” in International Conference on Latent Variable Analysis and Signal Separation . Springer, 2017, pp. 258–266
2017
Earlier work this paper cites.
A. Jansson, E. Humphrey, N. Montecchio, R. Bittner, A. Kumar, and T. Weyde, “Singing voice separation with deep U-Net convolutional networks,” in 18th International Society for Music Information Retrieval Conference , 2017, pp. 23–27
2017
Cited alongside, same era.
S. Uhlich, M. Porcu, F. Giron, M. Enenkl, T. Kemp, N. Takahashi, and Y. Mitsufuji, “Improving music source separation based on deep neural networks through data augmentation and network blending,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 261–265
2017
Cited alongside, same era.
H. Choi, J.-h. Lee, and K. Lee, “Singing voice separation using generative adversarial networks,,” in ML4Audio Workshop, 31st Conf. Neural Information Processing Systems (NIPS 2017) , 2017
2017
Cited alongside, same era.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 609–617
2017
G. Meseguer-Brocal and G. Peeters, “Conditioned-u-net: Introducing a control mechanism in the u-net for multiple source separations,” in 20th International Society for Music Information Retrieval Conference (ISMIR) , 2019
2019
Later among the works it cites.
S. Wisdom, J. R. Hershey, K. Wilson, J. Thorpe, M. Chinen, B. Patton, and R. A. Saurous, “Differentiable consistency constraints for improved deep speech enhancement,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 900–904
2019
Later among the works it cites.
2019
Later among the works it cites.
F.-R. Stöter, S. Uhlich, A. Liutkus, and Y. Mitsufuji, “Open-Unmix – A Reference Implementation for Music Source Separation,” Journal of Open Source Software , 2019. [Online]. Available: https://doi.org/10.21105/joss.01667
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 570–586
2018
Cited alongside, same era.
R. Gao, R. Feris, and K. Grauman, “Learning to separate object sounds by watching unlabeled video,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 35–53
2018
Cited alongside, same era.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 2, pp. 423–443, 2018
2018
Cited alongside, same era.
B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,” in Advances in Neural Information Processing Systems , 2018, pp. 7763–7774
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Tasnet: time-domain audio separation network for real-time, single-channel speech separation,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 696–700
2018
Cited alongside, same era.
D. Stoller, S. Ewert, and S. Dixon, “Adversarial semi-supervised audio source separation applied to singing voice extraction,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 2391–2395
2018
Cited alongside, same era.
——, “Objects that sound,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 435–451
2018
Cited alongside, same era.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 631–648
2018
Cited alongside, same era.
2019
Later among the works it cites.
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. Wilson, J. Le Roux, and J. R. Hershey, “Universal sound separation,” in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2019, pp. 175–179
2019
Later among the works it cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR – half-baked or well done?” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 626–630
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Parekh, S. Essid, A. Ozerov, N. Q. Duong, P. Pérez, and G. Richard, “Weakly supervised representation learning for audio-visual scene analysis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 416–428, 2019
2019
Later among the works it cites.
J. Y. Liu, Y. H. Yang, and S. K. Jeng, “Weakly-Supervised Visual Instrument-Playing Action Detection in Videos,” IEEE Transactions on Multimedia , vol. 21, no. 4, pp. 887–901, 4 2019
2019
Later among the works it cites.
R. Lu, Z. Duan, and C. Zhang, “Audio–visual deep clustering for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 11, pp. 1697–1712, 2019
2019
Later among the works it cites.
S. Parekh, A. Ozerov, S. Essid, N. Q. Duong, P. Pérez, and G. Richard, “Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision,” in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2019, pp. 268–272
2019
Later among the works it cites.
O. Slizovskaia, L. Kim, G. Haro, and E. Gomez, “End-to-end sound source separation conditioned on instrument labels,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 306–310
2019
Later among the works it cites.
K. Schulze-Forster, C. Doire, G. Richard, and R. Badeau, “Weakly informed audio source separation,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , 2019, pp. 268–272
2019
Later among the works it cites.
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms,” Proc. Interspeech 2019 , pp. 2350–2354, 2019
2019
Later among the works it cites.
A. Liutkus and F.-R. Stöter, “sigsep/norbert: First official norbert release,” Jul. 2019. [Online]. Available: https://doi.org/10.5281/zenodo.3269749
2019
Later among the works it cites.
R. Gao, T.-H. Oh, K. Grauman, and L. Torresani, “Listen to look: Action recognition by previewing audio,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 457–10 467
2020
Closest in time.
E. Tzinis, S. Wisdom, J. R. Hershey, A. Jansen, and D. P. Ellis, “Improving universal sound separation using sound classification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 96–100
2020
Closest in time.
F. Pishdadian, G. Wichern, and J. L. Roux, “Finding strength in weakness: Learning to separate sounds with weak supervision,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , pp. 2386–2399, 2020
2020
Closest in time.
J. F. Montesinos, O. Slizovskaia, and G. Haro, “Solos: A dataset for audio-visual music source separation and localization,” in Proceedings of the IEEE 22nd International Workshop on Multimedia Signal Processing , 2020
2020
Closest in time.
D. Michelsanti, Z.-H. Tan, S.-X. Zhang, Y. Xu, M. Yu, D. Yu, and J. Jensen, “An overview of deep-learning-based audio-visual speech enhancement and separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1368–1396, 2021
2021
Closest in time.