Fetching the paper…
Reading the bibliography…
The previous SpEx+ has yielded outstanding performance in speaker extraction and attracted much attention.
N. Mesgarani and E. F. Chang, “Selective cortical representation of attended speaker in multi-talker speech perception,” Nature , vol. 485, no. 7397, pp. 233–236, 2012
2012
Earlier work this paper cites.
2016
Earlier work this paper cites.
E. M. Kaya and M. Elhilali, “Modelling auditory attention,” Philosophical Transactions of the Royal Society B: Biological Sciences , vol. 372, no. 1714, p. 20160101, 2017
2017
Earlier work this paper cites.
K. Zmolikova, M. Delcroix, K. Kinoshita, T. Higuchi, A. Ogawa, and T. Nakatani, “Speaker-aware neural network based beamformer for speaker extraction in speech mixtures.” in Interspeech , 2017, pp. 2655–2659
2017
Earlier work this paper cites.
D. T. Toledano, M. P. Fernández-Gallego, and A. Lozano-Diez, “Multi-resolution speech analysis for automatic speech recognition using deep neural networks: Experiments on TIMIT,” PloS one , vol. 13, no. 10, p. e0205355, 2018
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
C. Xu, W. Rao, E. S. Chng, and H. Li, “Time-domain speaker extraction network,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 327–334
2019
Earlier work this paper cites.
K. Žmolíková, M. Delcroix, K. Kinoshita, T. Ochiai, T. Nakatani, L. Burget, and J. Černockỳ, “Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 4, pp. 800–814, 2019
2019
Earlier work this paper cites.
M. Delcroix, K. Zmolikova, T. Ochiai, K. Kinoshita, S. Araki, and T. Nakatani, “Compact network for speakerbeam target speaker extraction,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6965–6969
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Xu, W. Rao, E. S. Chng, and H. Li, “SpEx: Multi-scale time domain speaker extraction network,” IEEE/ACM transactions on audio, speech, and language processing , vol. 28, pp. 1370–1384, 2020
2020
Cited alongside, same era.
Q. Wang, I. L. Moreno, M. Saglam, K. Wilson, A. Chiao, R. Liu, Y. He, W. Li, J. Pelecanos, M. Nika et al. , “VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition,” Proc. Interspeech 2020 , pp. 2677–2681, 2020
2020
Cited alongside, same era.
M. Delcroix, T. Ochiai, K. Zmolikova, K. Kinoshita, N. Tawara, T. Nakatani, and S. Araki, “Improving speaker discrimination of target speech extraction with time-domain speakerbeam,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 691–695
2020
Cited alongside, same era.
S. He, H. Li, and X. Zhang, “Speakerfilter: Deep learning-based target speaker extraction using anchor speech,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 376–380
Y. Ju, W. Rao, X. Yan, Y. Fu, S. Lv, L. Cheng, Y. Wang, L. Xie, and S. Shang, “TEA-PSE: Tencent-ethereal-audio-lab personalized speech enhancement system for ICASSP 2022 DNS CHALLENGE,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 9291–9295
2022
Later among the works it cites.
C. Li, Y. Wang, F. Deng, Z. Zhang, X. Wang, and Z. Wang, “EAD-Conformer: a Conformer-Based Encoder-Attention-Decoder-Network for Multi-Task Audio Source Separation,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 521–525
2022
Later among the works it cites.
Y. Zhang, Z. Lv, H. Wu, S. Zhang, P. Hu, Z. Wu, H.-y. Lee, and H. Meng, “MFA-Conformer: Multi-scale Feature Aggregation Conformer for Automatic Speaker Verification,” Proc. Interspeech 2022 , pp. 306–310, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
M. Ge, C. Xu, L. Wang, E. S. Chng, J. Dang, and H. Li, “SpEx+: A Complete Time Domain Speaker Extraction Network,” Proc. Interspeech 2020 , pp. 1406–1410, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
D. Kim, J. Kim, and J. Kim, “Elastic exponential linear units for convolutional neural networks,” Neurocomputing , vol. 406, pp. 253–266, 2020
2020
Cited alongside, same era.
S. Chen, Y. Wu, Z. Chen, J. Wu, J. Li, T. Yoshioka, C. Wang, S. Liu, and M. Zhou, “Continuous speech separation with conformer,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5749–5753
2021
Cited alongside, same era.
W. Wang, C. Xu, M. Ge, and H. Li, “Neural Speaker Extraction with Speaker-Speech Cross-Attention Network.” in Interspeech , 2021, pp. 3535–3539
2021
Cited alongside, same era.
D. Min, D. B. Lee, E. Yang, and S. J. Hwang, “Meta-stylespeech: Multi-speaker adaptive text-to-speech generation,” in International Conference on Machine Learning . PMLR, 2021, pp. 7748–7759
2021
Cited alongside, same era.
J. Chen, Z. Wang, D. Tuo, Z. Wu, S. Kang, and H. Meng, “FullSubNet+: Channel Attention FullSubNet with Complex Spectrograms for Speech Enhancement,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7857–7861
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Lei, S. Yang, J. Cong, L. Xie, and D. Su, “Glow-WaveGAN 2: High-quality Zero-shot Text-to-speech Synthesis and Any-to-any Voice Conversion,” Proc. Interspeech 2022 , pp. 2563–2567, 2022
2022
Later among the works it cites.
J. Han, Y. Long, L. Burget, and J. Černockỳ, “DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation And Extraction,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7292–7296
2022
Later among the works it cites.
Z. Zhao, D. Yang, R. Gu, H. Zhang, and Y. Zou, “Target Confusion in End-to-end Speaker Extraction: Analysis and Approaches,” Proc. Interspeech 2022 , pp. 5333–5337, 2022
2022
Later among the works it cites.
Y. Ju, S. Zhang, W. Rao, Y. Wang, T. Yu, L. Xie, and S. Shang, “TEA-PSE 2.0: Sub-Band Network for Real-Time Personalized Speech Enhancement,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 472–479
2023
Closest in time.