Fetching the paper…
Reading the bibliography…
This paper presents Daft-Exprt, a multi-speaker acoustic model advancing the state-of-the-art for cross-speaker prosody transfer on any text.
S. King and V. Karaiskos, “The blizzard challenge 2013,” Blizzard Challenge Workshop, 2013
2013
Earlier work this paper cites.
M. I. C. Aarestrup, L. C. Jensen, and K. Fischer, “The Sound Makes the Greeting: Interpersonal Functions of Intonation in Human-Robot Interaction,” in AAAI Spring Symposium , 2015
2015
Earlier work this paper cites.
Y. Ganin and V. Lempitsky, “Unsupervised Domain Adaptation by Backpropagation,” in ICML , 2015
2015
Earlier work this paper cites.
ITU, “Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems (MUSHRA),” Tech. Rep., 2015
2015
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards End-to-end Speech Synthesis,” in INTERSPEECH , 2017
2017
Earlier work this paper cites.
V. Dumoulin, J. Shlens, and M. Kudlur, “A Learned Representation for Artistic Style,” in ICLR , 2017
2017
Earlier work this paper cites.
X. Huang and S. Belongie, “Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization,” in ICCV , 2017
2017
Earlier work this paper cites.
T. Kim, I. Song, and Y. Bengio, “Dynamic Layer Normalization for Adaptive Neural Acoustic Modeling in Speech Recognition,” in INTERSPEECH , 2017
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions,” in ICASSP , 2018
2018
Earlier work this paper cites.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous, “Towards End-to-end Prosody Transfer for Expressive Speech Synthesis with Tacotron,” in ICML , 2018
2018
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous, “Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-end Speech Synthesis,” in ICML , 2018
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. Courville, “FiLM: Visual Reasoning with a General Conditioning Layer,” in AAAI , 2018
2018
Earlier work this paper cites.
B. N. Oreshkin, P. Rodriguez, and A. Lacoste, “TADAM: Task Dependent Adaptive Metric for Improved Few-shot Learning,” in NeurIPS , 2018
2018
Earlier work this paper cites.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PLOS ONE , 2018
2018
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, “Neural Speech Synthesis with Transformer Network,” in AAAI , 2019
2019
Cited alongside, same era.
Y. Lee and T. Kim, “Robust and Fine-grained Prosody Control of End-to-end Speech Synthesis,” in ICASSP , 2019
2019
Cited alongside, same era.
Y. Bian, C. Chen, Y. Kang, and Z. Pan, “Multi-reference Tacotron by Intercross Training for Style Disentangling , Transfer and Control in Speech Synthesis,” in INTERSPEECH , 2019
2019
Cited alongside, same era.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning Latent Representations for Style Control and Transfer in End-to-end Speech Synthesis,” in ICASSP , 2019
E. Battenberg, S. Mariooryad, D. Stanton, R. Skerry-Ryan, M. Shannon, D. Kao, and T. Bagby, “Effective Use of Variational Embedding Capacity in Expressive End-to-end Speech Synthesis,” in ICLR , 2020
2020
Later among the works it cites.
J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu, “Non-attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling,” arXiv , 2020
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” in NeurIPS , 2020
2020
Later among the works it cites.
T. Li, X. Wang, Q. Xie, Z. Wang, and L. Xie, “Controllable Cross-speaker Emotion Transfer for End-to-end Speech Synthesis,” arXiv , 2021
2021
Closest in time.
X. An, F. K. Soong, and L. Xie, “Improving Performance of Seen and Unseen Speech Style Transfer in End-to-end Neural TTS,” in INTERSPEECH , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, P. Nguyen, and R. Pang, “Hierarchical Generative Modeling for Controllable Speech Synthesis,” in ICLR , 2019
2019
Cited alongside, same era.
Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Z. Chen, R. Skerry-Ryan, Y. Jia, A. Rosenberg, and B. Ramabhadran, “Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-language Voice Cloning,” in INTERSPEECH , 2019
2019
Cited alongside, same era.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-speech,” in INTERSPEECH , 2019
2019
Cited alongside, same era.
J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, R. Barra-Chicote, A. Moinet, and V. Aggarwal, “Towards Achieving Robust Universal Neural Vocoding,” in INTERSPEECH , 2019
2019
Cited alongside, same era.
O. Watts, G. E. Henter, J. Fong, and C. Valentini-Botinhao, “Where Do the Improvements Come From in Sequence-to-sequence Neural TTS?” in ISCA Speech Synthesis Workshop , 2019
2019
Cited alongside, same era.
S. Karlapati, A. Moinet, A. Joly, V. Klimkov, D. Sá\parez-Trigueros, and T. Drugman, “CopyCat: Many-to-many Fine-grained Prosody Transfer for Neural Text-to-speech,” in INTERSPEECH , 2020
2020
Cited alongside, same era.
R. Valle, J. Li, R. Prenger, and B. Catanzaro, “Mellotron: Multispeaker Expressive Voice Synthesis by Conditioning on Rhythm, Pitch and Global Style Tokens,” in ICASSP , 2020
2020
Cited alongside, same era.
2021
Closest in time.
R. Valle, K. Shih, R. Prenger, and B. Catanzaro, “Flowtron: an Autoregressive Flow-based Generative Network for Text-to-speech Synthesis,” in ICLR , 2021
2021
Closest in time.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and High-quality End-to-end Text to Speech,” in ICLR , 2021
2021
Closest in time.
A. Łań\parcucki, “FastPitch: Parallel Text-to-speech with Pitch Prediction,” in ICASSP , 2021
2021
Closest in time.
K. Lee, K. Park, and D. Kim, “STYLER: Style Factor Modeling with Rapidity and Robustness via Speech Decomposition for Expressive and Controllable Neural Text to Speech,” in INTERSPEECH , 2021
2021
Closest in time.
Z. Shang, Z. Huang, H. Zhang, P. Zhang, and Y. Yan, “Incorporating Cross-speaker Style Transfer for Multi-language Text-to-speech,” in INTERSPEECH , 2021
2021
Closest in time.
Y.-H. Chen, D.-Y. Wu, T.-H. Wu, and H.-Y. Lee, “AGAIN-VC: A One-shot Voice Conversion Using Activation Guidance and Adaptive Instance Normalization,” in ICASSP , 2021
2021
Closest in time.
M. Chen, X. Tan, B. Li, Y. Liu, T. Qin, S. Zhao, and T.-Y. Liu, “AdaSpeech: Adaptive Text to Speech for Custom Voice,” in ICLR , 2021
2021
Closest in time.
D. Min, D. B. Lee, E. Yang, and S. J. Hwang, “Meta-StyleSpeech: Multi-speaker Adaptive Text-to-speech Generation,” in ICML , 2021
2021
Closest in time.