Fetching the paper…
Reading the bibliography…
This paper describes Asteroid, the PyTorch-based audio source separation toolkit for researchers.
D. Griffin and J. Lim, “Signal estimation from modified short-time Fourier transform,”
1984
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ) — a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,”
2006
Earlier work this paper cites.
K. Nakadai, H. G. Okuno, H. Nakajima, Y. Hasegawa, and H. Tsujino, “An open source software system for robot audition HARK and its evaluation,” in
2008
Earlier work this paper cites.
B. Schuller, A. Lehmann, F. Weninger, F. Eyben, and G. Rigoll, “Blind enhancement of the rhythmic and harmonic sections by nmf: Does it help?” in
2009
Earlier work this paper cites.
D. Gunawan and D. Sen, “Iterative phase estimation for the synthesis of separated sources from single-channel mixtures,”
2010
Earlier work this paper cites.
S. van der Walt, S. C. Colbert, and G. Varoquaux, “The NumPy array: A structure for efficient numerical computation,”
2011
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek
2011
Earlier work this paper cites.
F. Grondin, D. Létourneau, F. Ferland, V. Rousseau, and F. Michaud, “The ManyEars open framework,”
2013
Earlier work this paper cites.
N. Perraudin, P. Balazs, and P. Søndergaard, “A fast Griffin-Lim algorithm,” in
2013
Earlier work this paper cites.
Y. Salaün, E. Vincent, N. Bertin, N. Souviraà-Labastie, X. Jaureguiberry, D. T. Tran, and F. Bimbot, “The Flexible Audio Source Separation Toolbox Version 2.0,” ICASSP Show & Tell, 2014
2014
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: discriminative embeddings for segmentation and separation,” in
2016
Earlier work this paper cites.
Y. Isik, J. Le Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” in
2016
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in
2017
Earlier work this paper cites.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in
2017
Cited alongside, same era.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
2017
Cited alongside, same era.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “The MUSDB18 corpus for music separation,” 2017. [Online]. Available:
2017
Cited alongside, same era.
L. Drude and R. Haeb-Umbach, “Tight integration of spatial and spectral features for BSS with deep clustering embeddings,” in
2017
Cited alongside, same era.
E. Vincent, T. Virtanen, and S. Gannot,
2018
Cited alongside, same era.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR — half-baked or well done?” in
2019
Later among the works it cites.
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux, “WHAM!: extending speech separation to noisy environments,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
W. Falcon
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Luo and N. Mesgarani, “TasNet: Time-domain audio separation network for real-time, single-channel speech separation,” in
2018
Cited alongside, same era.
E. Manilow, P. Seetharaman, and B. Pardo, “The Northwestern University Source Separation Library,” in
2018
Cited alongside, same era.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with SincNet,” in
2018
Cited alongside, same era.
J. M. Martín-Doñas, A. M. Gomez, J. A. Gonzalez, and A. M. Peinado, “A deep learning loss function based on the perceptual evaluation of the speech quality,”
2018
Cited alongside, same era.
——, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
2019
Cited alongside, same era.
Z. Ni and M. I. Mandel, “Onssen: an open-source speech separation and enhancement library,”
2019
Cited alongside, same era.
F.-R. Stöter, S. Uhlich, A. Liutkus, and Y. Mitsufuji, “Open-Unmix - A Reference Implementation for Music Source Separation,”
2019
Cited alongside, same era.
N. Zeghidour and D. Grangier, “Wavesplit: End-to-end speech separation by speaker clustering,”
2020
Closest in time.
E. Tzinis, S. Venkataramani, Z. Wang, C. Subakan, and P. Smaragdis, “Two-step sound source separation: Training on learned latent targets,” in
2020
Closest in time.
Y. Luo, Z. Chen, and T. Yoshioka, “Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech separation,” in
2020
Closest in time.
J. Heitkaemper, D. Jakobeit, C. Boeddeker, L. Drude, and R. Haeb-Umbach, “Demystifying TasNet: A dissecting approach,” in
2020
Closest in time.
M. Pariente, S. Cornell, A. Deleforge, and E. Vincent, “Filterbank design for end-to-end speech separation,” in
2020
Closest in time.
D. Ditter and T. Gerkmann, “A multi-phase gammatone filterbank for speech separation via TasNet,” in
2020
Closest in time.
J. Cosentino, S. Cornell, M. Pariente, A. Deleforge, and E. Vincent, “Librimix,”
2020
Closest in time.
S. Wisdom, H. Erdogan, D. P. W. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. R. Hershey, “What’s all the fuss about free universal sound separation data?”
2020
Closest in time.
C. K. A. Reddy, E. Beyrami, H. Dubey, V. Gopal, R. Cheng
2020
Closest in time.
M. Maciejewski, G. Wichern, E. McQuinn, and J. Le Roux, “WHAMR!: Noisy and reverberant single-channel speech separation,” in
2020
Closest in time.