Fetching the paper…
Reading the bibliography…
Supervised deep learning approaches to underdetermined audio source separation achieve state-of-the-art performance but require a dataset of mixtures along with their corresponding isolated source signals.
N. Levinson, “The Wiener (root mean square) error criterion in filter design and prediction,” J. Mathematics and Physics , vol. 25, no. 1-4, pp. 261–278, 1946
1946
Earlier work this paper cites.
B. L. Welch, “The generalization of ‘Student’s’ problem when several different population variances are involved,” Biometrika , vol. 34, no. 1-2, pp. 28–35, 1947
1947
Earlier work this paper cites.
J. Durbin, “The fitting of time-series models,” Revue de l’Institut International de Statistique , pp. 233–244, 1960
1960
Earlier work this paper cites.
H. Levene, “Robust tests for equality of variances,” Contributions to probability and statistics. Essays in honor of Harold Hotelling , pp. 279–292, 1961
1961
Earlier work this paper cites.
G. Fant, Acoustic theory of speech production . Walter de Gruyter, 1970, no. 2
1970
Earlier work this paper cites.
F. Itakura, “Line spectrum representation of linear predictor coefficients of speech signals,” The Journal of the Acoustical Society of America , vol. 57, no. S1, pp. S35–S35, 1975
1975
Earlier work this paper cites.
F. Soong and B. Juang, “Line spectrum pair (LSP) and speech data compression,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing , vol. 9. IEEE, 1984, pp. 37–40
1984
Earlier work this paper cites.
P. Kabal and R. P. Ramachandran, “The computation of line spectral frequencies using Chebyshev polynomials,” IEEE/ACM Trans. on Audio, Speech, and Language Processing , vol. 34, no. 6, pp. 1419–1426, 1986
1986
Earlier work this paper cites.
X. Serra and J. Smith, “Spectral modeling synthesis: A sound analysis/synthesis system based on a deterministic plus stochastic decomposition,” Computer Music Journal , vol. 14, no. 4, pp. 12–24, 1990
1990
Earlier work this paper cites.
D. H. Klatt and L. C. Klatt, “Analysis, synthesis, and perception of voice quality variations among female and male talkers,” The Journal of the Acoustical Society of America , vol. 87, no. 2, pp. 820–857, 1990
1990
Earlier work this paper cites.
J. Laroche, Y. Stylianou, and E. Moulines, “HNM: A simple, efficient harmonic + noise model for speech,” in Proc. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics . IEEE, 1993, pp. 169–172
1993
Earlier work this paper cites.
G. Richard and C. d’Alessandro, “Analysis/synthesis and modification of the speech aperiodic component,” Speech Communication , vol. 19, no. 3, pp. 221–244, 1996
1996
Earlier work this paper cites.
D. D. Lee and H. S. Seung, “Learning the parts of objects by non-negative matrix factorization,” Nature , vol. 401, no. 6755, pp. 788–791, 1999
1999
Earlier work this paper cites.
P. Horowitz and W. Hill, The art of electronics . Cambridge university press Cambridge, 2002
2002
Earlier work this paper cites.
E. Chew and X. Wu, “Separating voices in polyphonic music: A contig mapping approach,” in Int. Symp. on Computer Music Modeling and Retrieval . Springer, 2004, pp. 1–20
2004
Earlier work this paper cites.
W. C. Chu, Speech coding algorithms: foundation and evolution of standardized coders . John Wiley & Sons, 2004
2004
Earlier work this paper cites.
2006
Earlier work this paper cites.
A. Klapuri, “Multiple fundamental frequency estimation by summing harmonic amplitudes.” in Proc. Int. Soc. Music Inf. Retrieval Conf. , 2006, pp. 216–221
2006
Earlier work this paper cites.
D. S. Moore and S. Kirkland, The basic practice of statistics . WH Freeman New York, 2007, vol. 2
2007
Earlier work this paper cites.
T. Virtanen, “Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria,” IEEE/ACM Trans. on Audio, Speech, and Language Processing , vol. 15, no. 3, pp. 1066–1074, 2007
2007
Earlier work this paper cites.
I. V. McLoughlin, “Line spectral pairs,” Signal processing , vol. 88, no. 3, pp. 448–467, 2008
2008
Earlier work this paper cites.
T. Heittola, A. Klapuri, and T. Virtanen, “Musical instrument recognition in polyphonic audio using source-filter model for sound separation.” in Proc. Int. Soc. Music Inf. Retrieval Conf. , 2009
2009
Earlier work this paper cites.
L. Rabiner and R. Schafer, Theory and applications of digital speech processing . Prentice Hall Press, 2010
2010
Earlier work this paper cites.
R. Hennequin, B. David, and R. Badeau, “Score informed audio source separation using a parametric model of non-negative spectrogram,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2011, pp. 45–48
2011
Earlier work this paper cites.
J.-L. Durrieu, B. David, and G. Richard, “A musically motivated mid-level representation for pitch estimation and musical audio source separation,” IEEE J. Selected Topics in Signal Processing , vol. 5, no. 6, pp. 1180–1191, 2011
2011
Cited alongside, same era.
J. O. Smith, Spectral Audio Signal Processing: Frequency sampling method . W3K Publishing, 2011. [Online]. Available: https://ccrma.stanford.edu/~jos/sasp/Frequency_Sampling_Method_FIR.html
2011
Cited alongside, same era.
J. O. Smith, Spectral Audio Signal Processing: Generalized Window Method . W3K Publishing, 2011. [Online]. Available: https://ccrma.stanford.edu/~jos/sasp/Generalized_Window_Method.html
2011
Cited alongside, same era.
S. Ewert and M. Müller, “Using score-informed constraints for NMF-based source separation,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2012, pp. 129–132
2012
Cited alongside, same era.
J. Engel, C. Gu, A. Roberts et al. , “DDSP: Differentiable digital signal processing,” in Proc. Int. Conf. Learning Representations , 2019
2019
Later among the works it cites.
G. Meseguer-Brocal and G. Peeters, “Conditioned-U-Net: Introducing a control mechanism in the U-Net for multiple source separations,” in Proc. Int. Soc. Music Inf. Retrieval Conf. , 2019
2019
Later among the works it cites.
L. Drude, D. Hasenklever, and R. Haeb-Umbach, “Unsupervised training of a deep clustering model for multichannel blind source separation,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2019, pp. 695–699
2019
Later among the works it cites.
E. Tzinis, S. Venkataramani, and P. Smaragdis, “Unsupervised deep clustering for source separation: Direct learning from mixtures using spatial information,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2019, pp. 81–85
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Degottex, P. Lanchantin, A. Roebel, and X. Rodet, “Mixed source model and its adapted vocal tract filter estimate for voice transformation and synthesis,” Speech Communication , vol. 55, no. 2, pp. 278–294, 2013
2013
Cited alongside, same era.
D. Erro, I. Sainz, E. Navas, and I. Hernaez, “Harmonics plus noise model based vocoder for statistical parametric speech synthesis,” IEEE J. Selected Topics in Signal Processing , vol. 8, no. 2, pp. 184–194, 2013
2013
Cited alongside, same era.
A. L. Maas, A. Y. Hannun, A. Y. Ng et al. , “Rectifier nonlinearities improve neural network acoustic models,” in Proc. Int. Conf. Machine Learning , vol. 30, no. 1. Citeseer, 2013, p. 3
2013
Cited alongside, same era.
K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder–decoder approaches,” Syntax, Semantics and Structure in Statistical Translation , p. 103, 2014
2014
Cited alongside, same era.
2015
Cited alongside, same era.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2016, pp. 31–35
2016
Cited alongside, same era.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “Tensorflow: A system for large-scale machine learning,” in 12th symposium on operating systems design and implementation) , 2016, pp. 265–283
2016
Cited alongside, same era.
A. McLeod and M. Steedman, “HMM-based voice separation of MIDI performance,” J. New Music Research , vol. 45, no. 1, pp. 17–26, 2016
2016
Cited alongside, same era.
S. Leglaive, L. Girin, and R. Horaud, “Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2019, pp. 101–105
2019
Later among the works it cites.
X. Wang, S. Takaki, and J. Yamagishi, “Neural source-filter-based waveform model for statistical parametric speech synthesis,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2019, pp. 5916–5920
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR–half-baked or well done?” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2019, pp. 626–630
2019
Later among the works it cites.
E. Demirel, S. Ahlbäck, and S. Dixon, “Automatic lyrics transcription using dilated convolutional neural networks with self-attention,” in Int. Joint Conf. on Neural Networks . IEEE, 2020, pp. 1–8
2020
Later among the works it cites.
2020
Later among the works it cites.
V. Narayanaswamy, J. J. Thiagarajan, R. Anirudh, and A. Spanias, “Unsupervised audio source separation using generative priors,” in Proc. Interspeech , 2020, pp. 2657–2661
2020
Later among the works it cites.
P. Seetharaman, G. Wichern, J. Le Roux, and B. Pardo, “Bootstrapping unsupervised deep music separation from primitive auditory grouping principles,” Workshop on Self-supervision in Audio and Speech at the 37th Int. Conf. Machine Learning , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
H. Cuesta, B. McFee, and E. Gómez, “Multiple f0 estimation in vocal ensembles using convolutional neural networks,” in Proc. Int. Soc. Music Inf. Retrieval Conf. , 2020, pp. 302–309
2020
Later among the works it cites.
D. Petermann, P. Chandna, H. Cuesta, J. Bonada, and E. Gomez, “Deep learning based source separation applied to choir ensembles,” in Proc. Int. Soc. Music Inf. Retrieval Conf. , 2020, pp. 733–739
2020
Later among the works it cites.
M. Gover and P. Depalle, “Score-informed source separation of choral music,” in Proc. Int. Soc. Music Inf. Retrieval Conf. , 2020, pp. 231–239
2020
Later among the works it cites.
T. Nakamura and H. Kameoka, “Harmonic-temporal factor decomposition for unsupervised monaural separation of harmonic sounds,” IEEE/ACM Trans. on Audio, Speech, and Language Processing , vol. 29, pp. 68–82, 2020
2020
Later among the works it cites.
A. Rao and P. K. Ghosh, “SFNet: A computationally efficient source filter model based neural speech synthesis,” IEEE Signal Processing Letters , vol. 27, pp. 1170–1174, 2020
2020
Later among the works it cites.
J. Engel, R. Swavely, L. H. Hantrakul, A. Roberts, and C. Hawthorne, “Self-supervised pitch detection by inverse audio synthesis,” Workshop on Self-supervision in Audio and Speech at the 37th Int. Conf. on Machine Learning , 2020
2020
Later among the works it cites.
W. Zhang, Z. Chen, and F. Yin, “Multi-pitch estimation of polyphonic music based on pseudo two-dimensional spectrum,” IEEE/ACM Trans. on Audio, Speech, and Language Processing , vol. 28, pp. 2095–2108, 2020
2020
Later among the works it cites.
2021
Later among the works it cites.
Y.-N. Hung, G. Wichern, and J. Le Roux, “Transcription is all you need: Learning to separate musical mixtures with score as supervision,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 2021, pp. 46–50
2021
Later among the works it cites.
V. Monga, Y. Li, and Y. C. Eldar, “Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing,” IEEE Signal Processing Magazine , vol. 38, no. 2, pp. 18–44, 2021
2021
Later among the works it cites.