Fetching the paper…
Reading the bibliography…
In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker.
“Generating sequences with recurrent neural networks,”
Alex Graves, · 2013
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Earlier work this paper cites.
“Emotional end-to-end neural speech synthesizer,”
Younggun Lee, Azam Rabiee, and Soo-Young Lee, · 2017
Earlier work this paper cites.
“Fully character-level neural machine translation without explicit segmentation,”
Jason Lee, Kyunghyun Cho, and Thomas Hofmann, · 2017
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“Emphatic speech generation with conditioned input layer and bidirectional lstms for expressive speech synthesis,”
Runnan Li, Zhiyong Wu, Yuchen Huang, Jia Jia, Helen Meng, and Lianhong Cai, · 2018
Earlier work this paper cites.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A Saurous, · 2018
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A Saurous, · 2018
Cited alongside, same era.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Cited alongside, same era.
“Multi-speaker emotional acoustic modeling for cnn-based speech synthesis,”
Heejin Choi, Sangjun Park, Jinuk Park, and Minsoo Hahn, · 2019
Cited alongside, same era.
“Location-relative attention mechanisms for robust long-form speech synthesis,”
Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, and Tom Bagby, · 2020
Later among the works it cites.
“Controllable emotion transfer for end-to-end speech synthesis,”
Tao Li, Shan Yang, Liumeng Xue, and Lei Xie, · 2021
Closest in time.
“Expressive tts training with frame and style reconstruction loss,”
Rui Liu, Berrak Sisman, Guang lai Gao, and Haizhou Li, · 2021
Closest in time.
“Cross-speaker style transfer with prosody bottleneck in neural speech synthesis,”
Shifeng Pan and Lei He, · 2021
Closest in time.
“Incorporating cross-speaker style transfer for multi-language text-to-speech,”
Zengqiang Shang, Zhihua Huang, Haozhe Zhang, Pengyuan Zhang, and Yonghong Yan, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yanyao Bian, Changbin Chen, Yongguo Kang, and Zhenglin Pan, · 2019
Cited alongside, same era.
“Multi-reference neural tts stylization with adversarial cycle consistency,”
Matt Whitehill, Shuang Ma, Daniel McDuff, and Yale Song, · 2019
Cited alongside, same era.
“Learning latent representations for style control and transfer in end-to-end speech synthesis,”
Ya-Jie Zhang, Shifeng Pan, Lei He, and Zhen-Hua Ling, · 2019
Cited alongside, same era.
“Copycat: Many-to-many fine-grained prosody transfer for neural text-to-speech,”
Sri Karlapati, Alexis Moinet, Arnaud Joly, Viacheslav Klimkov, Daniel Sáez-Trigueros, and Thomas Drugman, · 2020
Cited alongside, same era.
“Controllable cross-speaker emotion transfer for end-to-end speech synthesis,”
Tao Li, Xinsheng Wang, Qicong Xie, Zhichao Wang, and Lei Xie, · 2021
Closest in time.
“Durian: Duration informed attention network for speech synthesis.,”
Chengzhu Yu, Heng Lu, Na Hu, Meng Yu, Chao Weng, Kun Xu, Peng Liu, Deyi Tuo, Shiyin Kang, Guangzhi Lei, et al., · 2031
Closest in time.