Fetching the paper…
Reading the bibliography…
Text-to-speech synthesis (TTS) has witnessed rapid progress in recent years, where neural methods became capable of producing audios with high naturalness.
Parallel neural text-to-speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao. 2019 · 1905
Earlier work this paper cites.
FastSpeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 1905
Earlier work this paper cites.
Shadowing
Sylvie Lambert. 1992 · 1992
Earlier work this paper cites.
The perfect voice
CARL Elliott. 2003 · 2003
Earlier work this paper cites.
A corpus study of the 3rd tone sandhi in standard chinese
Yiya Chen and Jiahong Yuan. 2007 · 2007
Earlier work this paper cites.
Towards incremental speech generation in dialogue systems
Gabriel Skantze and Anna Hjalmarsson. 2010 · 2010
Earlier work this paper cites.
Real-time incremental speech-to-speech translation of dialogs
Srinivas Bangalore, Vivek Kumar Rangarajan Sridhar, Prakash Kolan, Ladan Golipour, and Aura Jimenez. 2012 · 2012
Earlier work this paper cites.
INPRO_iSS: A component for just-in-time incremental speech synthesis
Timo Baumann and David Schlangen. 2012b · 2012
Earlier work this paper cites.
The INPROTK 2012 release
Timo Baumann and David Schlangen. 2012c · 2012
Earlier work this paper cites.
Combining incremental language generation and incremental speech synthesis for adaptive information presentation
Hendrik Buschmeier, Timo Baumann, Benjamin Dosch, Stefan Kopp, and David Schlangen. 2012 · 2012
Earlier work this paper cites.
Decision tree usage for incremental parametric speech synthesis
Timo Baumann. 2014a · 2014
Cited alongside, same era.
3 rd tone sandhi in standard chinese: A corpus approach
Jiahong Yuan and Yiya Chen. 2014 · 2014
Cited alongside, same era.
HMM training strategy for incremental speech synthesis
Maël Pouget, Thomas Hueber, Gérard Bailly, and Timo Baumann. 2015 · 2015
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Montreal Forced Aligner: Trainable text-speech alignment using Kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Cited alongside, same era.
Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention
Hideyuki Tachibana, Katsuya Uenoyama, and Shunsuke Aihara. 2018 · 2018
Later among the works it cites.
Incremental TTS for Japanese language
Tomoya Yanagita, Sakriani Sakti, and Satoshi Nakamura. 2018 · 2018
Later among the works it cites.
FloWaveNet: A generative flow for raw audio
Sungwon Kim, Sang-Gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon. 2019 · 2019
Closest in time.
Neural speech synthesis with Transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Closest in time.
STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework
Mingbo Ma, Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, Hua Wu, and Haifeng Wang. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan Ömer Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John L. Miller. 2017 · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous. 2017 · 2017
Cited alongside, same era.
Parallel WaveNet: Fast high-fidelity speech synthesis
Aaron Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et al. 2018 · 2018
Cited alongside, same era.
ClariNet: Parallel wave generation in end-to-end text-to-speech
Wei Ping, Kainan Peng, and Jitong Chen. 2018 · 2018
Cited alongside, same era.
Natural TTS synthesis by conditioning WaveNet on MEL spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, Rif Saurous, Yannis Agiomvrgiannakis, and Yonghui Wu. 2018 · 2018
Cited alongside, same era.
Partial representations improve the prosody of incremental speech synthesis
Timo Baumann. 2014b
Cited in the paper.
Evaluating prosodic processing for incremental speech synthesis
Timo Baumann and David Schlangen. 2012a
Cited in the paper.
Sashi Novitasari, Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2019 · 2019
Closest in time.
Waveglow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2019 · 2019
Closest in time.
Neural iTTS: Toward Synthesizing Speech in Real-time with End-to-end Neural Text-to-Speech Framework
Tomoya Yanagita, Sakriani Sakti, and Satoshi Nakamura. 2019 · 2019
Closest in time.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. 2020 · 2020
Closest in time.