Fetching the paper…
Reading the bibliography…
We propose a unified model for three inter-related tasks: 1) to \textit{separate} individual sound sources from a mixed music audio, 2) to \textit{transcribe} each sound source to MIDI notes, and 3) to\textit{ synthesize} new pieces based on the timbre of separated sources.
J. F. Woodruff, B. Pardo, and R. B. Dannenberg, “Remixing stereo music with score-informed source separation.” in ISMIR . Citeseer, 2006, pp. 314–319
2006
Earlier work this paper cites.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research , vol. 12, pp. 2825–2830, 2011
2011
Earlier work this paper cites.
C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. Ellis, and C. C. Raffel, “Mir_eval: A transparent implementation of common mir metrics,” in In Proceedings of the 15th International Society for Music Information Retrieval Conference, ISMIR . Citeseer, 2014
2014
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 31–35
2016
Earlier work this paper cites.
M. Miron, J. Janer Mestres, and E. Gómez Gutiérrez, “Monaural score-informed source separation for classical music using convolutional neural networks,” in 18th International Society for Music Information Retrieval Conference (ISMIR) , 2017
2017
Earlier work this paper cites.
S. Ewert and M. B. Sandler, “Structured dropout for weak label and multi-instance learning and its application to score-informed source separation,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 2277–2281
2017
Earlier work this paper cites.
A. Jansson, E. Humphrey, N. Montecchio, R. Bittner, A. Kumar, and T. Weyde, “Singing voice separation with deep u-net convolutional networks,” in International Society for Music Information Retrieval Conference (ISMIR) , 2017, pp. 23–27
2017
Earlier work this paper cites.
T. Yoshioka, H. Erdogan, Z. Chen, and F. Alleva, “Multi-microphone neural speech separation for far-field multi-talker speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5739–5743
2018
Earlier work this paper cites.
T. Yoshioka, H. Erdogan, Z. Chen, and F. Alleva, “Multi-microphone neural speech separation for far-field multi-talker speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5739–5743
2018
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 570–586
2018
Earlier work this paper cites.
N. Takahashi, N. Goswami, and Y. Mitsufuji, “Mmdenselstm: An efficient combination of convolutional and recurrent neural networks for audio source separation,” in 2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC) . IEEE, 2018, pp. 106–110
2018
Cited alongside, same era.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Alternative objective functions for deep clustering,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 686–690
2018
Cited alongside, same era.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 4, pp. 787–796, 2018
2018
Cited alongside, same era.
R. Kumar, Y. Luo, and N. Mesgarani, “Music source activity detection and separation using deep attractor network.” in INTERSPEECH , 2018, pp. 347–351
2018
Cited alongside, same era.
E. Manilow, G. Wichern, and J. Le Roux, “Hierarchical musical instrument separation,” in 21st International Society for Music Information Retrieval Conference (ISMIR) , 2020
2020
Later among the works it cites.
M. Gover, “Score-informed source separation of choral music,” in 21st International Society for Music Information Retrieval Conference (ISMIR) , 2020
2020
Later among the works it cites.
G. Meseguer-Brocal and G. Peeters, “Content based singing voice source separation via strong conditioning using aligned phonemes,” in 21st International Society for Music Information Retrieval Conference (ISMIR) , 2020
2020
Later among the works it cites.
C. Gan, D. Huang, H. Zhao, J. B. Tenenbaum, and A. Torralba, “Music gesture for visual sound separation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 478–10 487
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,” IEEE Transactions on Multimedia , vol. 21, no. 2, pp. 522–535, 2018
2018
Cited alongside, same era.
J. H. Lee, H.-S. Choi, and K. Lee, “Audio query-based music source separation,” in Proceedings of the 20th International Society for Music Information Retrieval Conference (ISMIR) , 2019
2019
Cited alongside, same era.
P. Seetharaman, G. Wichern, S. Venkataramani, and J. Le Roux, “Class-conditional embeddings for music source separation,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 301–305
2019
Cited alongside, same era.
G. Meseguer-Brocal and G. Peeters, “Conditioned-u-net: Introducing a control mechanism in the u-net for multiple source separations,” in Proceedings of the 20th International Society for Music Information Retrieval Conference (ISMIR) , 2019
2019
Cited alongside, same era.
D. Samuel, A. Ganeshan, and J. Naradowsky, “Meta-learning extractors for music source separation,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 816–820
2020
Cited alongside, same era.
E. Manilow, P. Seetharaman, and B. Pardo, “Simultaneous separation and transcription of mixtures with multiple polyphonic and percussive instruments,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 771–775
2020
Later among the works it cites.
Y.-N. Hung and A. Lerch, “Multitask learning for instrument activation aware music source separation,” in 21st International Society for Music Information Retrieval Conference (ISMIR) , 2020
2020
Later among the works it cites.
K. Tanaka, T. Nakatsuka, R. Nishikimi, K. Yoshii, and S. Morishima, “Multi-instrument music transcription based on deep spherical clustering of spectrograms and pitchgrams,” in 21st International Society for Music Information Retrieval Conference (ISMIR) , 2020
2020
Later among the works it cites.
C.-B. Jeon, H.-S. Choi, and K. Lee, “Exploring aligned lyrics-informed singing voice separation,” in 21st International Society for Music Information Retrieval Conference (ISMIR) , 2020
2020
Later among the works it cites.
A. Sharma, P. Kumar, V. Maddukuri, N. Madamshetti, K. Kishore, S. S. S. Kavuru, B. Raman, and P. P. Roy, “Fast griffin lim based waveform generation strategy for text-to-speech synthesis,” Multimedia Tools and Applications , vol. 79, no. 41, pp. 30 205–30 233, 2020
2020
Later among the works it cites.