Fetching the paper…
Reading the bibliography…
Single-channel speech separation has recently made great progress thanks to learned filterbanks as used in ConvTasNet.
“Parametric coding of speech spectra,”
J. L. Flanagan, · 1980
Earlier work this paper cites.
“Adam: a method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR,”
Felix Weninger, Hakan Erdogan, Shinji Watanabe, Emmanuel Vincent, Jonathan Le Roux, John R. Hershey, and Björn Schuller, · 2015
Earlier work this paper cites.
“Deep clustering: discriminative embeddings for segmentation and separation,”
John R. Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Single-channel multi-speaker separation using deep clustering,”
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R. Hershey, · 2016
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
D. Yu, M. Kolbæk, Z. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbaek, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Deep attractor network for single-microphone speaker separation,”
Zhuo Chen, Yi Luo, and Nima Mesgarani, · 2017
Cited alongside, same era.
“Alternative objective functions for deep clustering,”
Z. Wang, J. L. Roux, and J. R. Hershey, · 2018
Cited alongside, same era.
“Tasnet: Time-domain audio separation network for real-time, single-channel speech separation,”
Yi Luo and Nima Mesgarani, · 2018
Cited alongside, same era.
“Audlet filter banks: a versatile analysis/synthesis framework using auditory frequency scales,”
Thibaud Necciari, Nicki Holighaus, Peter Balazs, Zdeněk Průša, Piotr Majdak, and Olivier Derrien, · 2018
Cited alongside, same era.
“Speaker recognition from raw waveform with sincnet,”
M. Ravanelli and Y. Bengio, · 2018
Cited alongside, same era.
“Deep attention gated dilated temporal convolutional networks with intra-parallel convolutional modules for end-to-end monaural speech separation,”
“Improved speech separation with time-and-frequency cross-domain joint embedding and clustering,”
Gene-Ping Yang, Chao-I Tuan, Hung-Yi Lee, and Lin shan Lee, · 2019
Closest in time.
“A multi-phase gammatone filterbank for speech separation via tasnet,” 2019
David Ditter and Timo Gerkmann, · 2019
Closest in time.
“WHAM!: extending speech separation to noisy environments,”
Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, and Jonathan Le Roux, · 2019
Closest in time.
“On Learning Interpretable CNNs with Parametric Modulated Kernel-Based Filters,”
Erfan Loweimi, Peter Bell, and Steve Renals, · 2019
Closest in time.
“Sdr – half-baked or well done?,”
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Closest in time.
“On the variance of the adaptive learning rate and beyond,” 2019
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ziqiang Shi, Huibin Lin, Liu Liu, Rujie Liu, Jiqing Han, and Anyan Shi, · 2019
Cited alongside, same era.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Cited alongside, same era.
“Scripts to generate the wsj0 hipster ambient mixtures dataset,” http://wham.whisper.ai/
Cited in the paper.
Closest in time.
“Lookahead optimizer: k steps forward, 1 step back,” 2019
Michael R. Zhang, James Lucas, Geoffrey Hinton, and Jimmy Ba, · 2019
Closest in time.