Fetching the paper…
Reading the bibliography…
We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another.
R. Bellman and R. Kalaba, “On adaptive control processes,”
1959
Earlier work this paper cites.
T. Drugman and A. Alwan, “Joint robust voicing detection and pitch estimation based on residual harmonics,” in
1976
Earlier work this paper cites.
A. Ljolje and M. Riley, “Automatic segmentation and labeling of speech,” in
1991
Earlier work this paper cites.
R. B. ITU-R, “1534-1,“method for the subjective assessment of intermediate quality levels of coding systems (mushra)”,”
2003
Earlier work this paper cites.
R. A. Clark, M. Podsiadlo, M. Fraser, C. Mayo, and S. King, “Statistical analysis of the blizzard challenge 2007 listening test results,” in
2007
Earlier work this paper cites.
W. Chu and A. Alwan, “Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,” in
2009
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”
2013
Earlier work this paper cites.
Mozilla, “Common voice,” 2013. [Online]. Available: https://voice.mozilla.org/en/datasets
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Cited alongside, same era.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in
2016
Cited alongside, same era.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2wav: End-to-end speech synthesis,” in
2017
Cited alongside, same era.
A. Gibiansky, S. Arik, G. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep voice 2: Multi-speaker neural text-to-speech,” in
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. Lee and T. Kim, “Robust and fine-grained prosody control of end-to-end speech synthesis,”
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
2018
Cited alongside, same era.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in
2018
Cited alongside, same era.
2018
Later among the works it cites.
S. Liu, J. Zhong, L. Sun, X. Wu, X. Liu, and H. Meng, “Voice conversion across arbitrary speakers based on a single target-speaker utterance,” in
2018
Later among the works it cites.
N. Prateek, M. Lajszczak, R. Barra-Chicote, T. Drugman, J. Lorenzo-Trueba, T. Merritt, S. Ronanki, and T. Wood, “In other news: A bi-style text-to-speech model for synthesizing newscaster voice with limited data,”
2019
Closest in time.
J. Latorre, J. Lachowicz, J. Lorenzo-Trueba, T. Merritt, T. Drugman, S. Ronanki, and V. Klimkov, “Effect of data reduction on sequence-to-sequence neural tts,” in
2019
Closest in time.