Fetching the paper…
Reading the bibliography…
Recent research on the time-domain audio separation networks (TasNets) has brought great success to speech separation.
“Some experiments on the recognition of speech, with one and with two ears,”
E Colin Cherry, · 1953
Earlier work this paper cites.
“Continuous speech recognition (csr-i) wall street journal (wsj0) news, complete. linguistic data consortium, philadelphia (1993),”
J Garofalo, D David Graff, D Paul, and D Pallett, · 1993
Earlier work this paper cites.
“Learning to forget: Continual prediction with lstm,”
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins, · 1999
Earlier work this paper cites.
“Gradient flow in recurrent nets: the difficulty of learning long-term dependencies,” 2001
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, et al., · 2001
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Experiments on parallel training of deep neural network using model averaging,”
Hang Su and Haoyu Chen, · 2015
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering,”
Kai Chen and Qiang Huo, · 2016
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, Jesper Jensen, Morten Kolbaek, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Automatic differentiation in PyTorch,”
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, · 2017
Cited alongside, same era.
“Tasnet: time-domain audio separation network for real-time, single-channel speech separation,”
Yi Luo and Nima Mesgarani, · 2018
Cited alongside, same era.
“An empirical evaluation of generic convolutional and recurrent networks for sequence modeling,”
Shaojie Bai, J Zico Kolter, and Vladlen Koltun, · 2018
Cited alongside, same era.
“Wave-u-net: A multi-scale neural network for end-to-end audio source separation,”
Daniel Stoller, Sebastian Ewert, and Simon Dixon, · 2018
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
“T-gsa: Transformer with gaussian-weighted self-attention for speech enhancement,”
Jaeyoung Kim, Mostafa El-Khamy, and Jungwon Lee, · 2019
Later among the works it cites.
“Very deep self-attention networks for end-to-end speech recognition,”
Ngoc-Quan Pham, Thai-Son Nguyen, Jan Niehues, Markus Müller, Sebastian Stüker, and Alexander Waibel, · 2019
Later among the works it cites.
“Multi-stride self-attention for speech recognition.,”
Kyu J Han, Jing Huang, Yun Tang, Xiaodong He, and Bowen Zhou, · 2019
Later among the works it cites.
“R-transformer: Recurrent neural network enhanced transformer,”
Zhiwei Wang, Yao Ma, Zitao Liu, and Jiliang Tang, · 2019
Later among the works it cites.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Sharp nearby, fuzzy far away: How neural language models use context,”
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky, · 2018
Cited alongside, same era.
“Light gated recurrent units for speech recognition,”
Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, and Yoshua Bengio, · 2018
Cited alongside, same era.
“End-to-end speech separation with unfolded iterative phase reconstruction,”
Zhong-Qiu Wang, Jonathan Le Roux, DeLiang Wang, and John R Hershey, · 2018
Cited alongside, same era.
“Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,”
Yi Luo, Zhuo Chen, and Takuya Yoshioka, · 2019
Cited alongside, same era.
“Divide and conquer: A deep casa approach to talker-independent monaural speaker separation,”
Yuzhou Liu and DeLiang Wang, · 2019
Cited alongside, same era.
“Face landmark-based speaker-independent audio-visual speech enhancement in multi-talker environments,”
Giovanni Morrone, Sonia Bergamaschi, Luca Pasa, Luciano Fadiga, Vadim Tikhanoff, and Leonardo Badino, · 2019
Cited alongside, same era.
Later among the works it cites.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Later among the works it cites.
“Fasnet: Low-latency adaptive beamforming for multi-microphone audio processing,”
Yi Luo, Enea Ceolini, Cong Han, Shih-Chii Liu, and Nima Mesgarani, · 2019
Later among the works it cites.
“Improved speech separation with time-and-frequency cross-domain joint embedding and clustering,”
Gene-Ping Yang, Chao-I Tuan, Hung-Yi Lee, and Lin-shan Lee, · 2019
Later among the works it cites.
“Sdr–half-baked or well done?,”
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey, · 2019
Later among the works it cites.
“Audio-visual recognition of overlapped speech for the lrs2 dataset,”
Jianwei Yu, Shi-Xiong Zhang, Jian Wu, Shahram Ghorbani, Bo Wu, Shiyin Kang, Shansong Liu, Xunying Liu, Helen Meng, and Dong Yu, · 2020
Later among the works it cites.
Jingjing Chen, Qirong Mao, and Dong Liu, · 2020
Later among the works it cites.
“Simplified self-attention for transformer-based end-to-end speech recognition,”
Haoneng Luo, Shiliang Zhang, Ming Lei, and Lei Xie, · 2020
Later among the works it cites.
“Furcanext: End-to-end monaural speech separation with dynamic gated dilated temporal convolutional networks,”
Liwen Zhang, Ziqiang Shi, Jiqing Han, Anyan Shi, and Ding Ma, · 2020
Later among the works it cites.
“Mixup-breakdown: a consistency training method for improving generalization of speech separation models,”
Max WY Lam, Jun Wang, Dan Su, and Dong Yu, · 2020
Later among the works it cites.