Fetching the paper…
Reading the bibliography…
In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings.
“Noise-robust dynamic time warping using plca features,”
Brian King, Paris Smaragdis, and Gautham J Mysore, · 1976
Earlier work this paper cites.
“Multi-style training for robust isolated-word speech recognition,”
Richard Lippmann, Edward Martin, and D Paul, · 1987
Earlier work this paper cites.
“Mel-cepstral distance measure for objective speech quality assessment,”
R Kubichek, · 1993
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
Mike Schuster and Kuldip K Paliwal, · 1997
Earlier work this paper cites.
Advances in network and acoustic echo cancellation
Jacob Benesty, Tomas Gänsler, Dennis R Morgan, M Mohan Sondhi, and Steven L Gay, · 2001
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ) - a new method for speech quality assessment of telephone networks and codecs,”
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, · 2001
Earlier work this paper cites.
Acoustic echo and noise control: A practical approach
Eberhard Hänsler and Gerhard Schmidt, · 2005
Earlier work this paper cites.
“Toward accurate dynamic time warping in linear time and space,”
Stan Salvador and Philip Chan, · 2007
Earlier work this paper cites.
“Direct speech-to-speech translation with a sequence-to-sequence model,”
Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu, · 2008
Earlier work this paper cites.
A perspective on stereophonic acoustic echo cancellation
Jacob Benesty, Constantin Paleologu, Tomas Gänsler, and Silviu Ciochină, · 2011
Earlier work this paper cites.
“An unsupervised approach to cochannel speech separation,”
Ke Hu and DeLiang Wang, · 2012
Earlier work this paper cites.
“Text-informed audio source separation using nonnegative matrix partial co-factorization,”
Luc Le Magoarou, Alexey Ozerov, and Ngoc QK Duong, · 2013
Earlier work this paper cites.
“Generating sequences with recurrent neural networks,”
Alex Graves, · 2013
Earlier work this paper cites.
“Acoustic echo control,”
Gerald Enzner, Herbert Buchner, Alexis Favrot, and Fabian Kuech, · 2014
Earlier work this paper cites.
“Deep learning for monaural speech separation,”
Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, and Paris Smaragdis, · 2014
Earlier work this paper cites.
“Speech separation of a target speaker based on deep neural networks,”
Jun Du, Yanhui Tu, Yong Xu, Lirong Dai, and Chin-Hui Lee, · 2014
Earlier work this paper cites.
“One-against-all weighted dynamic time warping for language-independent and speaker-dependent speech recognition in adverse conditions,”
Xianglilan Zhang, Jiping Sun, and Zhigang Luo, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Text-informed speech enhancement with deep neural networks,”
Keisuke Kinoshita, Marc Delcroix, Atsunori Ogawa, and Tomohiro Nakatani, · 2015
Cited alongside, same era.
“Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,”
Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux, · 2015
Cited alongside, same era.
“Convolutional LSTM network: A machine learning approach for precipitation nowcasting,”
SHI Xingjian, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo, · 2015
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Cited alongside, same era.
“Single channel target speaker extraction and recognition with speaker beam,”
Marc Delcroix, Katerina Zmolikova, Keisuke Kinoshita, Atsunori Ogawa, and Tomohiro Nakatani, · 2018
Later among the works it cites.
“Multi-source syntactic neural machine translation,”
Anna Currey and Kenneth Heafield, · 2018
Later among the works it cites.
“Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Later among the works it cites.
“Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A Saurous, · 2018
Later among the works it cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, Yonghui Wu, et al., · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“LibriSpeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Cited alongside, same era.
“Superseded-CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2016
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al., · 2016
Cited alongside, same era.
“Tensorflow: A system for large-scale machine learning,”
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al., · 2016
Cited alongside, same era.
“Deliberation networks: Sequence generation beyond one-pass decoding,”
Yingce Xia, Fei Tian, Lijun Wu, Jianxin Lin, Tao Qin, Nenghai Yu, and Tie-Yan Liu, · 2017
Cited alongside, same era.
“Attention strategies for multi-source sequence-to-sequence learning,”
Jindřich Libovickỳ and Jindřich Helcl, · 2017
Cited alongside, same era.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Later among the works it cites.
“Deep multitask acoustic echo cancellation.,”
Amin Fazel, Mostafa El-Khamy, and Jungwon Lee, · 2019
Later among the works it cites.
“Deep neural network based regression approach for acoustic echo cancellation,”
Qinhui Lei, Hang Chen, Junfeng Hou, Liang Chen, and Lirong Dai, · 2019
Later among the works it cites.
“VoiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,”
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John R Hershey, Rif A Saurous, Ron J Weiss, Ye Jia, and Ignacio Lopez Moreno, · 2019
Later among the works it cites.
“Parrotron: An end-to-end speech-to-speech conversion model and its applications to hearing-impaired speech and speech separation,”
Fadi Biadsy, Ron J. Weiss, Pedro J. Moreno, Dimitri Kanvesky, and Ye Jia, · 2019
Later among the works it cites.
“Robust sequence-to-sequence acoustic modeling with stepwise monotonic attention for neural TTS,”
Mutian He, Yan Deng, and Lei He, · 2019
Later among the works it cites.
“Attention-based WaveNet autoencoder for universal voice conversion,”
Adam Polyak and Lior Wolf, · 2019
Later among the works it cites.
“LibriTTS: A corpus derived from LibriSpeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Later among the works it cites.
“Lingvo: A modular and scalable framework for sequence-to-sequence modeling,”
Jonathan Shen, Patrick Nguyen, Yonghui Wu, Zhifeng Chen, Mia X Chen, Ye Jia, Anjuli Kannan, Tara Sainath, Yuan Cao, Chung-Cheng Chiu, et al., · 2019
Later among the works it cites.
“Deliberation model based two-pass end-to-end speech recognition,”
Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar, · 2020
Closest in time.
“CAD-AEC: Context-aware deep acoustic echo cancellation,”
Amin Fazel, Mostafa El-Khamy, and Jungwon Lee, · 2020
Closest in time.
“Improved noisy student training for automatic speech recognition,”
Daniel S Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V Le, · 2020
Closest in time.