Fetching the paper…
Reading the bibliography…
In this work, we extend our previously proposed offline SpatialNet for long-term streaming multichannel speech enhancement in both static and moving speaker scenarios.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
S. Winter, W. Kellermann, H. Sawada, and S. Makino, “MAP-Based Underdetermined Blind Source Separation of Convolutive Mixtures by Hierarchical Clustering and
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, and B.-H. Juang, “Speech dereverberation based on variance-normalized delayed linear prediction,”
2010
Earlier work this paper cites.
E. A. Lehmann and A. M. Johansson, “Diffuse Reverberation Model for Efficient Image-Source Simulation of Room Impulse Responses,”
2010
Earlier work this paper cites.
T. Yoshioka and T. Nakatani, “Generalization of Multi-Channel Linear Prediction Methods for Blind MIMO Impulse Response Shortening,”
2012
Earlier work this paper cites.
J. Heymann, L. Drude, A. Chinaev, and R. Haeb-Umbach, “BLSTM supported GEV beamformer front-end for the 3RD CHiME challenge,” in
2015
Earlier work this paper cites.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines,” in
2015
Earlier work this paper cites.
D. P. Kingma and J. L. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in
2016
Earlier work this paper cites.
J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,”
2016
Earlier work this paper cites.
D. Yu, M. Kolbaek, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in
2017
Earlier work this paper cites.
S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, “A Consolidated Perspective on Multimicrophone Speech Enhancement and Source Separation,”
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Å. Kaiser, and I. Polosukhin, “Attention is All you Need,” in
2017
Cited alongside, same era.
C. Boeddecker, J. Heitkaemper, J. Schmalenstroeer, L. Drude, J. Heymann, and R. Haeb-Umbach, “Front-end processing for the CHiME-5 dinner party scenario,” in
2018
Cited alongside, same era.
M. Taseska and E. A. P. Habets, “Blind source separation of moving sources using sparsity-based source detection and tracking,”
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation,”
2019
Cited alongside, same era.
X. Li, L. Girin, S. Gannot, and R. Horaud, “Multichannel Online Dereverberation based on Spectral Magnitude Inverse Filtering,”
2019
Cited alongside, same era.
H. Chen, Y. Yang, F. Dang, and P. Zhang, “Beam-Guided TasNet: An Iterative Speech Separation Framework with Multi-Channel Output,” in
2022
Later among the works it cites.
A. Li, W. Liu, C. Zheng, and X. Li, “Embedding and Beamforming: All-Neural Causal Beamformer for Multichannel Speech Enhancement,” in
2022
Later among the works it cites.
Z.-Q. Wang, S. Cornell, S. Choi, Y. Lee, B.-Y. Kim, and S. Watanabe, “Tf-gridnet: Integrating full- and sub-band modeling for speech separation,”
2023
Later among the works it cites.
R. Gu, S.-X. Zhang, Y. Zou, and D. Yu, “Towards Unified All-Neural Beamforming for Time and Frequency Domain Speech Separation,”
2023
Later among the works it cites.
T. Ochiai, M. Delcroix, T. Nakatani, and S. Araki, “Mask-Based Neural Beamforming for Moving Speakers With Self-Attention-Based Tracking,”
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Nakatani and K. Kinoshita, “A Unified Convolutional Beamformer for Simultaneous Denoising and Dereverberation,”
2019
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in
2019
Cited alongside, same era.
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR – Half-baked or Well Done?” in
2019
Cited alongside, same era.
T. Ochiai, M. Delcroix, R. Ikeshita, K. Kinoshita, T. Nakatani, and S. Araki, “Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain Beamformer,” in
2020
Cited alongside, same era.
Y. Luo, Z. Chen, and T. Yoshioka, “Dual-Path RNN: Efficient Long Sequence Modeling for Time-Domain Single-Channel Speech Separation,” in
2020
Cited alongside, same era.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,” in
2020
Cited alongside, same era.
N. Moritz, T. Hori, and J. Le, “Streaming Automatic Speech Recognition with the Transformer Model,” in
2020
Cited alongside, same era.
K. Tesch and T. Gerkmann, “Insights Into Deep Non-Linear Filters for Improved Multi-Channel Speech Enhancement,”
2023
Later among the works it cites.
Y. Yang, C. Quan, and X. Li, “MCNET: Fuse Multiple Cues for Multichannel Speech Enhancement,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Gu and T. Dao, “Mamba: Linear-Time Sequence Modeling with Selective State Spaces,”
2023
Later among the works it cites.
Y. Sun, L. Dong, B. Patra, S. Ma, S. Huang, A. Benhaim, V. Chaudhary, X. Song, and F. Wei, “A Length-Extrapolatable Transformer,” in
2023
Later among the works it cites.
Y. Wang, A. Politis, and T. Virtanen, “Attention-driven multichannel speech enhancement in moving sound source scenarios,” in
2024
Closest in time.
C. Quan and X. Li, “SpatialNet: Extensively Learning Spatial Information for Multichannel Joint Speech Separation, Denoising and Dereverberation,”
2024
Closest in time.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “RoFormer: Enhanced transformer with Rotary Position Embedding,”
2024
Closest in time.