Fetching the paper…
Reading the bibliography…
Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,”
2006
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
2011
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z. H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in
2017
Earlier work this paper cites.
Y. Wu and K. He, “Group normalization,” in
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in
2018
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,”
2019
Earlier work this paper cites.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma
2019
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,”
2019
Earlier work this paper cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR — half-baked or well done?” in
2019
Earlier work this paper cites.
Y. Luo, Z. Chen, and T. Yoshioka, “Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech separation,” in
2020
Cited alongside, same era.
J. Chen, Q. Mao, and D. Liu, “Dual-Path Transformer Network: Direct context-aware modeling for end-to-end monaural speech separation,” in
2020
Cited alongside, same era.
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang
2020
Cited alongside, same era.
Y. Lu, Z. Li, D. He, Z. Sun, B. Dong
2020
Cited alongside, same era.
C. Li, J. Shi, W. Zhang, A. S. Subramanian, X. Chang
2021
Later among the works it cites.
N. Zeghidour and D. Grangier, “Wavesplit: End-to-end speech separation by speaker clustering,”
2021
Later among the works it cites.
L. Yang, W. Liu, and W. Wang, “TFPSNet: Time-frequency domain path scanning network for speech separation,” in
2022
Later among the works it cites.
T. Cord-Landwehr, C. Boeddeker, T. Von Neumann, C. Zorilă, R. Doddipatla
2022
Later among the works it cites.
J. Rixen and M. Renz, “QDPN - quasi-dual-path network for single-channel speech separation,” in
2022
Later among the works it cites.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,” in
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
M. Maciejewski, G. Wichern, E. McQuinn, and J. Le Roux, “WHAMR!: Noisy and reverberant single-channel speech separation,” in
2020
Cited alongside, same era.
C. K. Reddy, V. Gopal, R. Cutler, E. Beyrami, R. Cheng
2020
Cited alongside, same era.
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai
2021
Cited alongside, same era.
J. Droppo and O. Elibol, “Scaling laws for acoustic models,” in
2021
Cited alongside, same era.
Y.-J. Lu, S. Cornell, X. Chang, W. Zhang, C. Li
2022
Later among the works it cites.
Z.-Q. Wang, S. Cornell, S. Choi, Y. Lee, B.-Y. Kim
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey
2023
Later among the works it cites.
W. Zhang, K. Saijo, Z.-Q. Wang, S. Watanabe, and Y. Qian, “Toward universal speech enhancement for diverse input conditions,” in
2023
Later among the works it cites.
L. Liu, H. Guan, J. Ma, W. Dai, G. Wang
2023
Later among the works it cites.
S. Zhao, Y. Ma, C. Ni, C. Zhang, H. Wang
2024
Closest in time.
Y. Lee, S. Choi, B.-Y. Kim, Z.-Q. Wang, and S. Watanabe, “Boosting unknown-number speaker separation with transformer decoder-based attractor,” in
2024
Closest in time.