Fetching the paper…
Reading the bibliography…
We introduce Wavesplit, an end-to-end source separation system.
Y. Linde, A. Buzo, and R. Gray, “An algorithm for vector quantizer design,” IEEE Transactions on communications , vol. 28, no. 1, pp. 84–95, 1980
1980
Earlier work this paper cites.
D. Griffin and Jae Lim, “Signal estimation from modified short-time fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , 1984
1984
Earlier work this paper cites.
J. S. Garofolo, D. Graff, D. Paul, and D. S. Pallett, “CSR-I (WSJ0) complete,” Linguistic Data Consortium, Tech. Rep., 1993
1993
Earlier work this paper cites.
S. Araki, S. Makino, H. Sawada, and R. Mukai, “Underdetermined blind speech separation with directivity pattern based continuous mask and ica,” in 2004 12th European Signal Processing Conference . IEEE, 2004, pp. 1991–1994
1994
Earlier work this paper cites.
S. T. Roweis, “One microphone source separation,” in Advances in neural information processing systems , 2001, pp. 793–799
2001
Earlier work this paper cites.
O. Yilmaz and S. Rickard, “Blind separation of speech mixtures via time-frequency masking,” IEEE Transactions on signal processing , vol. 52, no. 7, pp. 1830–1847, 2004
2004
Earlier work this paper cites.
I. Mccowan, G. Lathoud, M. Lincoln, A. Lisowska, W. Post, D. Reidsma, and P. Wellner, “The ami meeting corpus,” in In: Proceedings Measuring Behavior 2005, 5th International Conference on Methods and Techniques in Behavioral Research. L.P.J.J. Noldus, F. Grieco, L.W.S. Loijens and P.H. Zimmerman (Eds.), Wageningen: Noldus Information Technology , 2005
2005
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Trans. Audio, Speech & Language Processing , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
J. Bernardes and D. Ayres-de Campos, “The persistent challenge of foetal heart rate monitoring,” Current Opinion in Obstetrics and Gynecology , vol. 22, no. 2, pp. 104–109, 2010
2010
Earlier work this paper cites.
E. Karvounis, M. Tsipouras, C. Papaloukas, D. Tsalikakis, K. Naka, and D. Fotiadis, “A non-invasive methodology for fetal monitoring during pregnancy,” Methods of information in medicine , vol. 49, no. 03, pp. 238–253, 2010
2010
Earlier work this paper cites.
D. Williamson, Discrete-time Signal Processing: An Algebraic Approach , ser. Advanced Textbooks in Control and Signal Processing. Springer, 2012
2012
Earlier work this paper cites.
P. C. Loizou, Speech enhancement: theory and practice . CRC press, 2013
2013
Earlier work this paper cites.
I. Silva, J. Behar, R. Sameni, T. Zhu, J. Oster, G. D. Clifford, and G. B. Moody, “Noninvasive fetal ecg: the physionet/computing in cardiology challenge 2013,” in Computing in Cardiology 2013 . IEEE, 2013, pp. 149–152
2013
Earlier work this paper cites.
I. Toumi, S. Caldarelli, and B. Torrésani, “A review of blind source separation in nmr spectroscopy,” Progress in nuclear magnetic resonance spectroscopy , vol. 81, pp. 37–64, 2014
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
C. Weng, D. Yu, M. L. Seltzer, and J. Droppo, “Deep neural networks for single-channel multi-talker speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, pp. 1670–1679, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015 , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR (Poster) , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. L. Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in ICASSP . IEEE, 2016, pp. 31–35
2016
Earlier work this paper cites.
N. Zeghidour, G. Synnaeve, N. Usunier, and E. Dupoux, “Joint learning of speaker and phonetic similarities with siamese networks,” in INTERSPEECH . ISCA, 2016, pp. 1295–1299
2016
Earlier work this paper cites.
F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” in 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” in SSW . ISCA, 2016, p. 125
2016
Cited alongside, same era.
Y. Isik, J. L. Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” in INTERSPEECH . ISCA, 2016, pp. 545–549
2016
Cited alongside, same era.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines,” in INTERSPEECH . ISCA, 2018, pp. 1561–1565
2018
Later among the works it cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in ICLR , 2018
2018
Later among the works it cites.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,” IEEE/ACM Trans. Audio, Speech & Language Processing , vol. 26, no. 4, pp. 787–796, 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Andreotti, J. Behar, S. Zaunseder, J. Oster, and G. D. Clifford, “An open-source framework for stress-testing non-invasive foetal ecg extraction algorithms,” Physiological measurement , vol. 37, no. 5, p. 627, 2016
2016
Cited alongside, same era.
D. Murray, L. Stankovic, and V. Stankovic, “An electrical load measurements dataset of united kingdom households from a two-year longitudinal study,” Scientific data , vol. 4, no. 1, pp. 1–12, 2017
2017
Cited alongside, same era.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “The MUSDB18 corpus for music separation,” Dec. 2017. [Online]. Available: https://doi.org/10.5281/zenodo.1117372
2017
Cited alongside, same era.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 241–245
2017
Cited alongside, same era.
M. Kolbaek, D. Yu, Z. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Trans. Audio, Speech & Language Processing , vol. 25, no. 10, pp. 1901–1913, 2017
2017
Cited alongside, same era.
J. R. Hershey, J. L. Roux, S. Watanabe, S. Wisdom, Z. Chen, and Y. Isik, “Novel deep architectures in speech processing,” in New Era for Robust Speech Recognition, Exploiting Deep Learning , 2017, pp. 135–164. [Online]. Available: https://doi.org/10.1007/978-3-319-64680-0_6
2017
Cited alongside, same era.
S. Uhlich, M. Porcu, F. Giron, M. Enenkl, T. Kemp, N. Takahashi, and Y. Mitsufuji, “Improving music source separation based on deep neural networks through data augmentation and network blending,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 261–265
2017
Cited alongside, same era.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in ICASSP . IEEE, 2017, pp. 246–250
2017
Cited alongside, same era.
2018
Later among the works it cites.
Z. Wang, J. L. Roux, and J. R. Hershey, “Alternative objective functions for deep clustering,” in ICASSP . IEEE, 2018, pp. 686–690
2018
Later among the works it cites.
Y. Luo and N. Mesgarani, “Tasnet: Time-domain audio separation network for real-time, single-channel speech separation,” in ICASSP . IEEE, 2018, pp. 696–700
2018
Later among the works it cites.
Y. Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio, Speech & Language Processing , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR - half-baked or well done?” in ICASSP . IEEE, 2019, pp. 626–630
2019
Later among the works it cites.
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux, “Wham!: Extending speech separation to noisy environments,” in Proc. Interspeech , Sep. 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Sablayrolles, M. Douze, C. Schmid, and H. Jégou, “Spreading vectors for similarity search,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 , 2019. [Online]. Available: https://openreview.net/forum?id=SkGuG2R5tm
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Closest in time.
L. Zhang, Z. Shi, J. Han, A. Shi, and D. Ma, “Furcanext: End-to-end monaural speech separation with dynamic gated dilated temporal convolutional networks,” in International Conference on Multimedia Modeling . Springer, 2020, pp. 653–665
2020
Closest in time.