Fetching the paper…
Reading the bibliography…
Recent studies in deep learning-based speech separation have proven the superiority of time-domain approaches to conventional time-frequency-based methods.
“Meeting transcription using virtual microphone arrays,”
Takuya Yoshioka, Zhuo Chen, Dimitrios Dimitriadis, William Hinthorn, Xuedong Huang, Andreas Stolcke, and Michael Zeng, · 1905
Earlier work this paper cites.
“Chunk-based decoder for neural machine translation,”
Shonosuke Ishiwatari, Jingtao Yao, Shujie Liu, Mu Li, Ming Zhou, Naoki Yoshinaga, Masaru Kitsuregawa, and Weijia Jia, · 1912
Earlier work this paper cites.
“Image method for efficiently simulating small-room acoustics,”
Jont B. Allen and David A. Berkley, · 1979
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Performance measurement in blind audio source separation,”
Emmanuel Vincent, Rémi Gribonval, and Cédric Févotte, · 2006
Earlier work this paper cites.
“Generating sensor signals in isotropic noise fields,”
Emanuël AP. Habets and Sharon Gannot, · 2007
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“A hierarchical neural autoencoder for paragraphs and documents,”
Jiwei Li, Minh-Thang Luong, and Dan Jurafsky, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Hierarchical multiscale recurrent neural networks,”
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio, · 2016
Earlier work this paper cites.
“SampleRNN: An unconditional end-to-end neural audio generation model,”
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio, · 2016
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R. Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Cited alongside, same era.
“Single-channel multi-speaker separation using deep clustering,”
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R Hershey, · 2016
Cited alongside, same era.
“Chunk-based bi-scale decoder for neural machine translation,”
Hao Zhou, Zhaopeng Tu, Shujian Huang, Xiaohua Liu, Hang Li, and Jiajun Chen, · 2017
Cited alongside, same era.
“Dilated recurrent neural networks,”
Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark A. Hasegawa-Johnson, and Thomas S Huang, · 2017
Cited alongside, same era.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Cited alongside, same era.
“Speaker-independent speech separation with deep attractor network,”
Yi Luo, Zhuo Chen, and Nima Mesgarani, · 2018
Later among the works it cites.
“End-to-end speech separation with unfolded iterative phase reconstruction,”
Zhong-Qiu Wang, Jonathan Le Roux, DeLiang Wang, and John R. Hershey, · 2018
Later among the works it cites.
“Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Closest in time.
“A comprehensive study of speech separation: Spectrogram vs waveform separation,”
Fahimeh Bahmaninezhad, Jian Wu, Rongzhi Gu, Shi-Xiong Zhang, Yong Xu, Meng Yu, and Dong Yu, · 2019
Closest in time.
“SDR–half-baked or well done?,”
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R. Hershey, · 2019
Closest in time.
“FaSNet: Low-latency adaptive beamforming for multi-microphone audio processing,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“TasNet: time-domain audio separation network for real-time, single-channel speech separation,”
Yi Luo and Nima Mesgarani, · 2018
Cited alongside, same era.
“Wave-U-Net: A multi-scale neural network for end-to-end audio source separation,”
Daniel Stoller, Sebastian Ewert, and Simon Dixon, · 2018
Cited alongside, same era.
“End-to-end source separation with adaptive front-ends,”
Shrikant Venkataramani, Jonah Casebeer, and Paris Smaragdis, · 2018
Cited alongside, same era.
“End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,”
Szu-Wei Fu, Tao-Wei Wang, Yu Tsao, Xugang Lu, and Hisashi Kawai, · 2018
Cited alongside, same era.
“Independently recurrent neural network (indrnn): Building a longer and deeper rnn,”
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao, · 2018
Cited alongside, same era.
“An empirical evaluation of generic convolutional and recurrent networks for sequence modeling,”
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun, · 2018
Cited alongside, same era.
“Real-time single-channel dereverberation and separation with time-domain audio separation network,”
Yi Luo and Nima Mesgarani, · 2018
Cited alongside, same era.
Yi Luo, Enea Ceolini, Cong Han, Shih-Chii Liu, and Nima Mesgarani, · 2019
Closest in time.
“End-to-end music source separation: Is it possible in the waveform domain?,”
Francesc Lluís, Jordi Pons, and Xavier Serra, · 2019
Closest in time.
“Universal sound separation,”
Ilya Kavalerov, Scott Wisdom, Hakan Erdogan, Brian Patton, Kevin Wilson, Jonathan Le Roux, and John R. Hershey, · 2019
Closest in time.
“Divide and conquer: A deep casa approach to talker-independent monaural speaker separation,”
Yuzhou Liu and DeLiang Wang, · 2019
Closest in time.
“Deep learning based phase reconstruction for speaker separation: A trigonometric perspective,”
Zhong-Qiu Wang, Ke Tan, and DeLiang Wang, · 2019
Closest in time.
“FurcaNeXt: End-to-end monaural speech separation with dynamic gated dilated temporal convolutional networks,”
Liwen Zhang, Ziqiang Shi, Jiqing Han, Anyan Shi, and Ding Ma, · 2020
Closest in time.