Fetching the paper…
Reading the bibliography…
Transformer based models have provided significant performance improvements in monaural speech separation.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
J. R. Hershey, Z. Chen, J. L. Roux, and S. Watanabe, · 2016
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“TasNet: time-domain audio separation network for real-time, single-channel speech separation,”
Y. Luo and N. Mesgarani, · 2018
Earlier work this paper cites.
“Alternative objective functions for deep clustering,”
Z.-Q. Wang, J. L. Roux, and J. R Hershey, · 2018
Earlier work this paper cites.
“Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Y. Luo and N. Mesgarani, · 2019
Earlier work this paper cites.
“Divide and conquer: A deep casa approach to talker-independent monaural speaker separation,”
Y. Liu and D. Wang, · 2019
Earlier work this paper cites.
“Wham!: Extending speech separation to noisy environments,”
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. L. Roux, · 2019
Earlier work this paper cites.
“Whamr!: Noisy and reverberant single-channel speech separation,”
M. Maciejewski, G. Wichern, E. McQuinn, and J. L. Roux, · 2019
Earlier work this paper cites.
“SDR – half-baked or well done?,”
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Cited alongside, same era.
“Deep learning based phase reconstruction for speaker separation: A trigonometric perspective,”
Z.-Q. Wang, K. Tan, and D. Wang, · 2019
Cited alongside, same era.
“Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech separation,”
Y. Luo, Z. Chen, and T. Yoshioka, · 2020
Cited alongside, same era.
“Voice separation with an unknown number of multiple speakers,”
E. Nachmani, Y. Adi, and L. Wolf, · 2020
Cited alongside, same era.
“Dual-Path Transformer Network: Direct context-aware modeling for end-to-end monaural speech separation,”
J. Chen, Q. Mao, and D. Liu, · 2020
Cited alongside, same era.
“Sudo rm -rf: Efficient networks for universal audio source separation,”
“Two-step sound source separation: Training on learned latent targets,”
E. Tzinis, S. Venkataramani, Z. Wang, C. Subakan, and P. Smaragdis, · 2020
Later among the works it cites.
“Multi-scale group transformer for long sequence modeling in speech separation,”
Y. Zhao, C. Luo, Z.-J. Zha, and W. Zeng, · 2020
Later among the works it cites.
“GLU variants improve transformer,”
N. Shazeer, · 2020
Later among the works it cites.
“Wavesplit: End-to-end speech separation by speaker clustering,”
N. Zeghidour and D. Grangier, · 2021
Later among the works it cites.
“Attention is all you need in speech separation,”
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, · 2021
Later among the works it cites.
“RoFormer: Enhanced transformer with rotary position embedding,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Tzinis, Z. Wang, and P. Smaragdis, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, · 2020
Cited alongside, same era.
“Filterbank design for end-to-end speech separation,”
M. Pariente, S. Cornell, A. Deleforge, and E. Vincent, · 2020
Cited alongside, same era.
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu, · 2021
Later among the works it cites.
“SepIt: Approaching a single channel speech separation bound,”
S. Lutati, E. Nachmani, and L. Wolf, · 2022
Later among the works it cites.
“Transformer quality in linear time,”
W. Hua, Z. Dai, H. Liu, and Q. Le, · 2022
Later among the works it cites.