Fetching the paper…
Reading the bibliography…
Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information.
E. C. Cherry, “Some experiments on the recognition of speech, with one and with two ears,” J. Acoust. Soc. Am. , vol. 25, no. 5, pp. 975–979, Sep. 1953
1953
Earlier work this paper cites.
J. Garofolo, D. Graff, D. Paul, and D. Pallett, “CSR-I (WSJ0) Complete LDC93S6A,” 1993
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neur. Comp. , vol. 9, no. 8, pp. 1735–1780, Nov. 1997
1997
Earlier work this paper cites.
S. Partan and P. Marler, “Communication goes multimodal,” Science , vol. 283, no. 5406, pp. 1272–1273, Feb. 1999
1999
Earlier work this paper cites.
A. F. Martin, M. A. Przybocki et al. , “Speaker recognition in a multi-speaker environment.” in Proc. Interspeech Conf. , 2001, pp. 787–790
2001
Earlier work this paper cites.
2005
Earlier work this paper cites.
D. E. King, “Dlib-ml: A machine learning toolkit,” The Journal of Machine Learning Research , vol. 10, pp. 1755–1758, Dec. 2009
2009
Earlier work this paper cites.
M. Cooke, J. R. Hershey, and S. J. Rennie, “Monaural speech separation and recognition challenge,” Computer Speech & Language , vol. 24, no. 1, pp. 1–15, Jan. 2010
2010
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2010, pp. 4214–4217
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 19, no. 4, pp. 788–798, May 2011
2011
Earlier work this paper cites.
E. Z. Golumbic, G. B. Cogan, C. E. Schroeder, and D. Poeppel, “Visual input enhances selective speech envelope tracking in auditory cortex at a “cocktail party”,” Journal of Neuroscience , vol. 33, no. 4, pp. 1417–1426, Jan. 2013
2013
Earlier work this paper cites.
J. Du, Y. Tu, Y. Xu, L. Dai, and C.-H. Lee, “Speech separation of a target speaker based on deep neural networks,” in Proc. Intl. Conf. on Signal Processing , Oct. 2014, pp. 473–477
2014
Earlier work this paper cites.
N. Harte and E. Gillen, “TCD-TIMIT: An audio-visual corpus of continuous speech,” IEEE Trans. Multimedia , vol. 17, no. 5, pp. 603–615, May 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: An ASR corpus based on public domain audio books,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2015, pp. 5206–5210
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. IEEE Intl. Conf. on Learn. Repr. (ICLR) , May 2015, pp. 1–15
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. L. Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2016, pp. 31–35
2016
Earlier work this paper cites.
J. Du, Y. Tu, L.-R. Dai, and C.-H. Lee, “A regression approach to single-channel speech separation via high-resolution deep neural networks,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 24, no. 8, pp. 1424–1437, Aug. 2016
2016
Earlier work this paper cites.
Y. Isik, J. L. Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” in Proc. Interspeech Conf. , Sep. 2016, pp. 545–549
2016
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z. H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2017, pp. 241–245
2017
Earlier work this paper cites.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Mar. 2017, pp. 246–250
2017
Earlier work this paper cites.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 25, no. 10, pp. 1901–1913, Oct. 2017
2017
Earlier work this paper cites.
F. Cole, D. Belanger, D. Krishnan, A. Sarna, I. Mosseri, and W. T. Freeman, “Synthesizing normalized faces from facial identity features,” in Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) , Jul. 2017, pp. 3703–3712
2017
Cited alongside, same era.
T. Stafylakis and G. Tzimiropoulos, “Combining residual networks with lstms for lipreading,” in Proc. Interspeech Conf. , Aug. 2017, pp. 3652–3656
2017
Cited alongside, same era.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 26, no. 10, pp. 1702–1726, Oct. 2018
2018
Cited alongside, same era.
M. Delcroix, K. Zmolikova, K. Kinoshita, A. Ogawa, and T. Nakatani, “Single channel target speaker extraction and recognition with speaker beam,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2018, pp. 5554–5558
2018
Cited alongside, same era.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR–half-baked or well done?” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , May 2019, pp. 626–630
2019
Later among the works it cites.
Y. Luo, Z. Chen, and T. Yoshioka, “Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech separation,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , May 2020, pp. 46–50
2020
Later among the works it cites.
J. Chen, Q. Mao, and D. Liu, “Dual-path transformer network: Direct context-aware modeling for end-to-end monaural speech separation,” in Proc. Interspeech Conf. , Oct. 2020, pp. 2642–2646
2020
Later among the works it cites.
C. Xu, W. Rao, E. S. Chng, and H. Li, “SpEx: Multi-scale time domain speaker extraction network,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 28, pp. 1370–1384, Apr. 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Wang, J. Chen, D. Su, L. Chen, M. Yu, Y. Qian, and D. Yu, “Deep extractor network for target speaker recovery from single channel speech mixtures,” in Proc. Interspeech Conf. , Sep. 2018, pp. 307–311
2018
Cited alongside, same era.
A. Gabbay, A. Shamir, and S. Peleg, “Visual speech enhancement,” in Proc. Interspeech Conf. , Sep. 2018, pp. 1170–1174
2018
Cited alongside, same era.
J.-C. Hou, S.-S. Wang, Y.-H. Lai, Y. Tsao, H.-W. Chang, and H.-M. Wang, “Audio-visual speech enhancement using multimodal deep convolutional neural networks,” IEEE Trans. on Emerging Topics in Computational Intelligence , vol. 2, no. 2, pp. 117–128, Apr. 2018
2018
Cited alongside, same era.
T. Afouras, J. S. Chung, and A. Zisserman, “The conversation: Deep audio-visual speech enhancement,” in Proc. Interspeech Conf. , Sep. 2018, pp. 3244–3248
2018
Cited alongside, same era.
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein, “Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,” ACM Trans. Graphics , vol. 37, no. 4, Jul. 2018
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “TasNet: Time-domain audio separation network for real-time, single-channel speech separation,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2018, pp. 696–700
2018
Cited alongside, same era.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2018, pp. 4879–4883
2018
Cited alongside, same era.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proc. of the European Conf. on Computer Vision (ECCV) , Sep. 2018
2018
Cited alongside, same era.
M. Delcroix, T. Ochiai, K. Zmolikova, K. Kinoshita, N. Tawara, T. Nakatani, and S. Araki, “Improving speaker discrimination of target speech extraction with time-domain speakerbeam,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , May 2020, pp. 691–695
2020
Later among the works it cites.
E. Ceolini, J. Hjortkjær, D. D. Wong, J. O’Sullivan, V. S. Raghavan, J. Herrero, A. D. Mehta, S.-C. Liu, and N. Mesgarani, “Brain-informed speech separation (BISS) for enhancement of target speaker in multitalker speech perception,” NeuroImage , vol. 223, p. 117282, Dec. 2020
2020
Later among the works it cites.
S.-Y. Chuang, Y. Tsao, C.-C. Lo, and H.-M. Wang, “Lite audio-visual speech enhancement,” in Proc. Interspeech Conf. , Oct. 2020, pp. 1131–1135
2020
Later among the works it cites.
M. Ge, C. Xu, L. Wang, E. S. Chng, J. Dang, and H. Li, “SpEx+: A complete time domain speaker extraction network,” in Proc. Interspeech Conf. , Oct. 2020, pp. 1406–1410
2020
Later among the works it cites.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” in Proc. Interspeech Conf. , Oct. 2020, pp. 3830–3834
2020
Later among the works it cites.
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “Voxceleb: Large-scale speaker verification in the wild,” Computer Speech & Language , vol. 60, p. 101027, Mar. 2020
2020
Later among the works it cites.
N. Zeghidour and D. Grangier, “Wavesplit: End-to-end speech separation by speaker clustering,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 2840–2849, Jul. 2021
2021
Later among the works it cites.
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Jun. 2021, pp. 21–25
2021
Later among the works it cites.
J. Byun and J. W. Shin, “Monaural speech separation using speaker embedding from preliminary separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 2753–2763, Aug. 2021
2021
Later among the works it cites.
M. Delcroix, K. Zmolikova, T. Ochiai, K. Kinoshita, and T. Nakatani, “Speaker activity driven neural speech extraction,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Jun. 2021, pp. 6099–6103
2021
Later among the works it cites.
D. Michelsanti, Z.-H. Tan, S.-X. Zhang, Y. Xu, M. Yu, D. Yu, and J. Jensen, “An overview of deep-learning-based audio-visual speech enhancement and separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 1368–1396, Mar. 2021
2021
Later among the works it cites.
Z. Aldeneh, A. P. Kumar, B.-J. Theobald, E. Marchi, S. Kajarekar, D. Naik, and A. H. Abdelaziz, “On the role of visual cues in audiovisual speech enhancement,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Jun. 2021, pp. 8423–8427
2021
Later among the works it cites.
S. S. Shetu, S. Chakrabarty, and E. A. P. Habets, “An empirical study of visual features for DNN based audio-visual speech enhancement in multi-talker environments,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Jun. 2021, pp. 8418–8422
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Zhao, D. Yang, R. Gu, H. Zhang, and Y. Zou, “Target confusion in end-to-end speaker extraction: Analysis and approaches,” in Proc. Interspeech Conf. , Sep. 2022, pp. 5333–5337
2022
Closest in time.
Z. Pan, M. Ge, and H. Li, “USEV: Universal speaker extraction with visual cue,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 30, pp. 3032–3045, Sep. 2022
2022
Closest in time.
Z.-Q. Wang, S. Cornell, S. Choi, Y. Lee, B.-Y. Kim, and S. Watanabe, “TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,” in Proc. IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , Jun. 2023, pp. 1–5
2023
Closest in time.