Fetching the paper…
Reading the bibliography…
Time-domain Transformer neural networks have proven their superiority in speech separation tasks.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in IEEE International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
J. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 31–35
2016
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 241–245
2017
Earlier work this paper cites.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multi-talker speech separation and tracing with permutation invariant training of deep recurrent neural networks,” IEEE Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 25, no. 10, pp. 1901–1913, 2017
2017
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Tasnet: time-domain audio separation network for real-time, single-channel speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 696–700
2018
Earlier work this paper cites.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaev, G. Venkatesh, and H. Wu, “Mixed precision training,” in IEEE International Conference on Learning Representations (ICLR) , 2018
2018
Earlier work this paper cites.
Y. Luo, N. Mesgarani, and Z. Chen, “Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,” IEEE Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Earlier work this paper cites.
A. Hannun, A. Lee, Q. Xu, and R. Collobert, “Sequence-to-sequence speech recognition with time-depth separable convolutions,” in IEEE Conference of the International Speech Communication Association (INTERSPEECH) , 2019, pp. 3785–3789
2019
Earlier work this paper cites.
J. Zhu, M. Hasegawa-Johnson, and L. Sari, “Identify speakers in cocktail parties with end-to-end attention,” in IEEE Conference of the International Speech Communication Association (INTERSPEECH) , 2020, pp. 3092–3096
2020
Earlier work this paper cites.
Q. Wang, I. Moreno, M. Saglam, K. Wilson, A. Chiao, R. Liu, Y. He, W. Li, J. Pelecanos, M. Nika, and A. Gruenstein, “Voicefilter-lite: Streaming targeted voice separation for on-device speech recognition,” in IEEE Conference of the International Speech Communication Association (INTERSPEECH) , 2020, pp. 2677–2681
2020
Earlier work this paper cites.
E. Tzinis, S. Venkataramani, Z. Wang, C. Subakan, and P. Smaragdis, “Two-step sound source separation: Training on learned latent targets,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 31–35
2020
Cited alongside, same era.
Y. Luo, Z. Chen, and T. Yoshioka, “Dual-path rnn: Efficient long sequence modeling for time-domain single-channel speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 46–50
2020
Cited alongside, same era.
J. Chen, Q. Mao, and D. Liu, “Dual-path transformer network: Direct context-aware modeling for end-to-end monaural speech separation,” in IEEE Conference of the International Speech Communication Association (INTERSPEECH) , 2020, pp. 46–50
2020
Cited alongside, same era.
S. Wang, B. Li, M. Khabsa, H. Fang, and H. Ma, “Linformer: Self-attention with linear complexity,” in IEEE Advances in Neural Information Processing Systems (NIPS) , 2020
2020
J. Luo, J. Wang, N. Cheng, G. Jiang, and J. Xiao, “Multi-quartznet: Multi-resolution convolution for speech recognition with multi-layer feature fusion,” in IEEE Spoken Language Technology Workshop (SLT) , 2021, pp. 82–88
2021
Later among the works it cites.
J. Luo, J. Wang, N. Cheng, E. Xiao, J. Xiao, G. Kucsko, P. O’Neill, J. Balam, S. Deng, A. Flores, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, and J. Li, “Cross-language transfer learning and domain adaptation for end-to-end automatic speech recognition,” in IEEE International Conference on Multimedia and Expo (ICME) , 2021, pp. 1–6
2021
Later among the works it cites.
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 21–25
2021
Later among the works it cites.
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, D. Belanger, L. Colwell, and A. Weller, “Rethinking attention with performers,” in IEEE International Conference on Learning Representations (ICLR) , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in IEEE Conference of the International Speech Communication Association (INTERSPEECH) , 2020, pp. 5036–5040
2020
Cited alongside, same era.
Z. Wu, Z. Liu, J. Lin, Y. Lin, and S. Han, “Lite transformer with long-short range attention,” in IEEE International Conference on Learning Representations (ICLR) , 2020
2020
Cited alongside, same era.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, R. Soricut, G. Research, and M. de Carvalho, “Albert: A lite bert for self-supervised learning of language representations,” in IEEE International Conference on Learning Representations (ICLR) , 2020
2020
Cited alongside, same era.
J. Luo, J. Wang, N. Cheng, and J. Xiao, “Unidirectional memory-self-attention transducer for online speech recognition,” in IEEE International Conference on Acoustics Speech and Signal Processing Proceedings (ICASSP) , 2021, pp. 910–914
2021
Cited alongside, same era.
X. Zhang, J. Qian, Y. Yu, Y. Sun, and W. Li, “Singer identification using deep timbre feature learning with knn-net,” in 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 3380–3384
2021
Cited alongside, same era.
S. Chen, Y. Wu, Z. Chen, J. Wu, J. Li, T. Yoshioka, C. Wang, S. Liu, and M. Zhou, “Continuous speech separation with conformer,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 5749–5753
2021
Cited alongside, same era.
2021
Later among the works it cites.
Y. Koizumi, S. Karita, S. Wisdom, H. Erdogan, J. Hershey, L. Jones, and M. Bacchiani, “Df-conformer: Integrated architecture of conv-tasnet and conformer using linear complexity self-attention for speech enhancement,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2021, pp. 161–165
2021
Later among the works it cites.
P.-H. Chi, P.-H. Chung, T.-H. Wu, C.-C. Hsieh, Y.-H. Chen, S.-W. Li, and H.-y. Lee, “Audio albert: A lite bert for self-supervised learning of audio representation,” in IEEE Spoken Language Technology Workshop (SLT) , 2021, pp. 344–350
2021
Later among the works it cites.
2021
Later among the works it cites.
N. Zeghidour and D. Grangier, “Wavesplit: End-to-end speech separation by speaker clustering,” IEEE Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 29, pp. 2840–2849, 2021
2021
Later among the works it cites.
D. Petermann, G. Wichern, Z.-Q. Wang, and J. Le Roux, “The cocktail fork problem: Three-stem audio separation for real-world soundtracks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 526–530
2022
Closest in time.
X. Zhang, J. Wang, N. Cheng, and J. Xiao, “Singer identification for metaverse with timbral and middle-level perceptual features,” in International Joint Conference on Neural Networks, (IJCNN) . IEEE, 2022, pp. 1–7
2022
Closest in time.