Fetching the paper…
Reading the bibliography…
The presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity.
ITU-T, Recommendation P.800: Methods for subjective determination of transmission quality , 1996
1996
Earlier work this paper cites.
——, Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs , 2001
2001
Earlier work this paper cites.
P. C. Loizou, Speech enhancement: theory and practice . CRC press, 2007
2007
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “DEMAND: A Collection of Multi-channel Recordings of Acoustic Noise in Diverse Environments,” June 2013. [Online]. Available: https://doi.org/10.5281/zenodo.1227121
2013
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 31–35
2016
Earlier work this paper cites.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 10, pp. 1901–1913, 2017
2017
Earlier work this paper cites.
K. Zmolikova, M. Delcroix, K. Kinoshita, T. Higuchi, A. Ogawa, and T. Nakatani, “Speaker-aware neural network based beamformer for speaker extraction in speech mixtures.” in Interspeech , 2017, pp. 2655–2659
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “VoxCeleb: A Large-Scale Speaker Identification Dataset,” in Proc. INTERSPEECH , 2017, pp. 2616–2620
2017
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Tasnet: Time-domain audio separation network for real-time, single-channel speech separation,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 696–700
2018
Earlier work this paper cites.
H. R. Muckenhirn, I. L. Moreno, J. Hershey, K. Wilson, P. Sridhar, Q. Wang, R. A. Saurous, R. Weiss, Y. Jia, and Z. Wu, “Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Earlier work this paper cites.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4879–4883
2018
Cited alongside, same era.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, z. Chen, P. Nguyen, R. Pang, I. Lopez Moreno, and Y. Wu, “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep Speaker Recognition,” in Proc. INTERSPEECH , 2018, pp. 1086–1090
2018
Cited alongside, same era.
ITU-T, Recommendation P.808: Subjective evaluation of speech quality with a crowdsourcing approach , 2018
2018
Cited alongside, same era.
J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 5891–5895
2019
Later among the works it cites.
J.-M. Valin, U. Isik, N. Phansalkar, R. Giri, K. Helwani, and A. Krishnaswamy, “A perceptually-motivated approach for low-complexity, real-time enhancement of fullband speech,” in Proc. INTERSPEECH . ISCA, 2020, pp. 2482–2486
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Mun, S. Choe, J. Huh, and J. S. Chung, “The sound of my voice: Speaker representation loss for target voice separation,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7289–7293
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
E. Tzinis, S. Venkataramani, and P. Smaragdis, “Unsupervised deep clustering for source separation: Direct learning from mixtures using spatial information,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 81–85
2019
Cited alongside, same era.
T. von Neumann, K. Kinoshita, M. Delcroix, S. Araki, T. Nakatani, and R. Haeb-Umbach, “All-neural online source separation, counting, and diarization for meeting analysis,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 91–95
2019
Cited alongside, same era.
R. Gu, L. Chen, S.-X. Zhang, J. Zheng, Y. Xu, M. Yu, D. Su, Y. Zou, and D. Yu, “Neural spatial filter: Target speaker speech separation assisted with directional information.” in Proc. INTERSPEECH , 2019, pp. 4290–4294
2019
Cited alongside, same era.
P. Wang, Z. Chen, X. Xiao, Z. Meng, T. Yoshioka, T. Zhou, L. Lu, and J. Li, “Speech separation using speaker inventory,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2019, pp. 230–236
2019
Cited alongside, same era.
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, “Fully supervised speaker diarization,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6301–6305
2019
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in Proceedings of the 36th International Conference on Machine Learning , 2019, pp. 5210–5219
2019
Cited alongside, same era.
2020
Later among the works it cites.
T. Li, Q. Lin, Y. Bao, and M. Li, “Atss-Net: Target speaker separation via attention-based neural network,” in Proc. INTERSPEECH , 2020, pp. 1411–1415
2020
Later among the works it cites.
Q. Wang, I. Lopez-Moreno, M. Saglam, K. Wilson, A. Chiao, R. Liu, Y. He, W. Li, J. Pelecanos, M. Nika, and A. Gruenstein, “Voicefilter-lite: Streaming targeted voice separation for on-device speech recognition,” in Proc. INTERSPEECH . ISCA, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Braun and I. Tashev, “Data augmentation and loss normalization for deep noise suppression,” in International Conference on Speech and Computer . Springer, 2020, pp. 79–86
2020
Later among the works it cites.
2020
Later among the works it cites.
U. Isik, R. Giri, N. Phansalkar, J.-M. Valin, K. Helwani, and A. Krishnaswamy, “PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Loss,” in Proc. INTERSPEECH , 2020, pp. 2487–2491
2020
Later among the works it cites.
S. Ding, Q. Wang, S.-Y. Chang, L. Wan, and I. Lopez Moreno, “Personal VAD: Speaker-conditioned voice activity detection,” in Proc. Odyssey 2020 The Speaker and Language Recognition Workshop , 2020, pp. 433–439
2020
Later among the works it cites.