Fetching the paper…
Reading the bibliography…
Existing CNN-based speech separation models face local receptive field limitations and cannot effectively capture long time dependencies.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
E. Grimm, R. Van Everdingen, and M. Schöpping, “Toward a recommendation for a european standard of peak and lkfs loudness levels,” SMPTE motion imaging journal , vol. 119, no. 3, pp. 28–34, 2010
2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 31–35
2016
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 241–245
2017
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in Conference of the International Speech Communication Association (Interspeech) , 2018
2018
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr–half-baked or well done?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 626–630
2019
Earlier work this paper cites.
E. Tzinis, Z. Wang, and P. Smaragdis, “Sudo rm-rf: Efficient networks for universal audio source separation,” in 2020 IEEE 30th International Workshop on Machine Learning for Signal Processing (MLSP) . IEEE, 2020, pp. 1–6
2020
Cited alongside, same era.
Y. Luo, Z. Chen, and T. Yoshioka, “Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 46–50
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
D. Petermann, G. Wichern, Z.-Q. Wang, and J. Le Roux, “The cocktail fork problem: Three-stem audio separation for real-world soundtracks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 526–530
2022
Later among the works it cites.
K. Li and Y. Luo, “On the design and training strategies for rnn-based online neural speech separation systems,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
Z.-Q. Wang, S. Cornell, S. Choi, Y. Lee, B.-Y. Kim, and S. Watanabe, “Tf-gridnet: Making time-frequency domain models great again for monaural speaker separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Hu, K. Li, W. Zhang, Y. Luo, J.-M. Lemercier, and T. Gerkmann, “Speech separation using an asynchronous fully recurrent convolutional neural network,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 34, pp. 22 509–22 522, 2021
2021
Cited alongside, same era.
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 21–25
2021
Cited alongside, same era.
A. Gu, K. Goel, and C. Re, “Efficiently modeling long sequences with structured state spaces,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré, “Combining recurrent, convolutional, and continuous-time models with linear state space layers,” in Advances in neural information processing systems , 2021, pp. 572–585
2021
Cited alongside, same era.
K. Li, R. Yang, and X. Hu, “An efficient encoder-decoder architecture with top-down attention for speech separation,” in The Eleventh International Conference on Learning Representations (ICLR) , 2022
2022
Cited alongside, same era.
K. Li, X. Hu, and Y. Luo, “On the Use of Deep Mask Estimation Module for Neural Source Separation Systems,” in Conference of the International Speech Communication Association (Interspeech) , 2022, pp. 5328–5332
2022
Cited alongside, same era.
L. Yang, W. Liu, and W. Wang, “Tfpsnet: Time-frequency domain path scanning network for speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6842–6846
2022
Cited alongside, same era.
2023
Later among the works it cites.
V. Sovrasov. (2023) ptflops: a flops counting tool for neural networks in pytorch framework. [Online]. Available: https://github.com/sovrasov/flops-counter.pytorch
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. Li, W. Sang, C. Zeng, R. Yang, G. Chen, and X. Hu. (2024) Librispace: A simulated audio toolkit for speech enhancement and separation. [Online]. Available: https://github.com/JusperLee/LibriSpace/
2024
Closest in time.