Fetching the paper…
Reading the bibliography…
Transformers have recently achieved state-of-the-art performance in speech separation.
“Language models are few-shot learners,”
T. Brown et al., · 1901
Earlier work this paper cites.
“Pruning algorithms-a survey,”
R. Reed, · 1993
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
G. Hinton, O. Vinyals, and J. Dean, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,” 2015,
D. P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
J. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, · 2016
Earlier work this paper cites.
“Pruning convolutional neural networks for resource efficient inference,”
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz, · 2017
Earlier work this paper cites.
“Quantized neural networks: Training neural networks with low precision weights and activations,”
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani et al., · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“Bitwise neural networks for efficient single-channel source separation,”
M. Kim and P. Smaragdis, · 2018
Earlier work this paper cites.
“TasNet: time-domain audio separation network for real-time, single-channel speech separation,”
Y. Luo and N. Mesgarani, · 2018
Earlier work this paper cites.
“Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Y. Luo and N. Mesgarani, · 2019
Earlier work this paper cites.
“Improving keyword spotting and language identification via neural architecture search at scale,”
H. Mazzawi et al., · 2019
Earlier work this paper cites.
“WHAM!: Extending speech separation to noisy environments,”
G. Wichern et al., · 2019
Earlier work this paper cites.
“SDR–half-baked or well done?,”
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Earlier work this paper cites.
“Deep learning based phase reconstruction for speaker separation: A trigonometric perspective,”
Z.-Q. Wang, K. Tan, and D. Wang, · 2019
Earlier work this paper cites.
“Divide and conquer: A deep CASA approach to talker-independent monaural speaker separation,”
Y. Liu and D. Wang, · 2019
Cited alongside, same era.
“Megatron-LM: Training multi-billion parameter language models using model parallelism,”
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
“Introducing the voiceprivacy initiative,”
N. Tomashenko et al., · 2020
Cited alongside, same era.
“PHASEN: A phase-and-harmonics-aware speech enhancement network,”
D. Yin, C. Luo, Z. Xiong, and W. Zeng, · 2020
Cited alongside, same era.
“DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,”
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
W.-N. Hsu et al., · 2021
Later among the works it cites.
“The energy and carbon footprint of training end-to-end speech recognizers,”
T. Parcollet and M. Ravanelli, · 2021
Later among the works it cites.
“Neural architecture search for LF-MMI trained time delay neural networks,”
S. Hu et al., · 2021
Later among the works it cites.
“Attention is all you need in speech separation,”
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, · 2021
Later among the works it cites.
“SpeechBrain: A general-purpose speech toolkit,”
M. Ravanelli et al., · 2021
Later among the works it cites.
“Wavesplit: End-to-end speech separation by speaker clustering,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Hu et al., · 2020
Cited alongside, same era.
“Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation,”
Y. Luo, Z. Chen, and T. Yoshioka, · 2020
Cited alongside, same era.
“SuDoRM-RF: Efficient networks for universal audio source separation,”
E. Tzinis, Z. Wang, and P. Smaragdis, · 2020
Cited alongside, same era.
“Linformer: Self-attention with linear complexity,”
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma, · 2020
Cited alongside, same era.
“Longformer: The long-document transformer,”
I. Beltagy, M. E. Peters, and A. Cohan, · 2020
Cited alongside, same era.
“Reformer: The efficient transformer,”
N. Kitaev, L. Kaiser, and A. Levskaya, · 2020
Cited alongside, same era.
“ESPnet-ST: All-in-one speech translation toolkit,”
H. Inaguma et al., · 2020
Cited alongside, same era.
N. Zeghidour and D. Grangier, · 2021
Later among the works it cites.
“PaLM: Scaling language modeling with pathways,”
A. Chowdhery et al., · 2022
Closest in time.
“SkiM: Skipping memory LSTM for low-latency real-time continuous speech separation,”
C. Li, L. Yang, W. Wang, and Y. Qian, · 2022
Closest in time.
“Neural architecture search for speech emotion recognition,”
X. Wu, S. Hu, Z. Wu, X. Liu, and H. Meng, · 2022
Closest in time.
“Efficient transformers: A survey,”
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, · 2022
Closest in time.
“Squeezeformer: An efficient transformer for automatic speech recognition,”
S. Kim et al., · 2022
Closest in time.
“SepIt: Approaching a single channel speech separation bound,”
S. Lutati, E. Nachmani, and L. Wolf, · 2022
Closest in time.
“TFPSNet: Time-frequency domain path scanning network for speech separation,”
L. Yang, W. Liu, and W. Wang, · 2022
Closest in time.
“Exploring self-attention mechanisms for speech separation,”
C. Subakan, M. Ravanelli, S. Cornell, F. Grondin, and M. Bronzi, · 2023
Closest in time.
“TF-GridNet: Integrating full- and sub-band modeling for speech separation,”
Z.-Q. Wang, S. Cornell, S. Choi, Y. Lee, B.-Y. Kim, and S. Watanabe, · 2023
Closest in time.
“MossFormer: Pushing the performance limit of monaural speech separation using gated single-head transformer with convolution-augmented joint self-attentions,”
S. Zhao and B. Ma, · 2023
Closest in time.