Fetching the paper…
Reading the bibliography…
We propose Cotatron, a transcription-guided speech encoder for speaker-independent linguistic representation.
R. L. Weide, “The cmu pronouncing dictionary,”
1998
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in
2016
Earlier work this paper cites.
A. van den Oord, O. Vinyals
2017
Earlier work this paper cites.
D. Ulyanov, A. Vedaldi, and V. S. Lempitsky, “Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis,” in
2017
Earlier work this paper cites.
V. Dumoulin, J. Shlens, and M. Kudlur, “A learned representation for artistic style,” in
2017
Earlier work this paper cites.
D. Ha, A. Dai, and Q. V. Le, “Hypernetworks,” in
2017
Earlier work this paper cites.
Y. Saito, Y. Ijima, K. Nishida, and S. Takamichi, “Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors,” in
2018
Earlier work this paper cites.
C.-c. Yeh, P.-c. Hsu, J.-c. Chou, H.-y. Lee, and L.-s. Lee, “Rhythm-flexible voice conversion without parallel data using cycle-gan over phoneme posteriorgram sequences,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu
2018
Earlier work this paper cites.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
H. Lu, Z. Wu, D. Dai, R. Li, S. Kang, J. Jia, and H. Meng, “One-shot voice conversion with global speaker embeddings,”
2019
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in
M. Zhang, X. Wang, F. Fang, H. Li, and J. Yamagishi, “Joint training framework for text-to-speech and voice conversion using multi-source tacotron and wavenet,”
2019
Later among the works it cites.
J. Yamagishi, C. Veaux, K. MacDonald
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Zen, R. Clark, R. J. Weiss, V. Dang, Y. Jia, Y. Wu, Y. Zhang, and Z. Chen, “Libritts: A corpus derived from librispeech for text-to-speech,”
2019
Later among the works it cites.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “Melgan: Generative adversarial networks for conditional waveform synthesis,” in
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Ding and R. Gutierrez-Osuna, “Group latent embedding for vector quantized variational autoencoder in non-parallel voice conversion,”
2019
Cited alongside, same era.
A. T. Liu, P.-c. Hsu, and H.-y. Lee, “Unsupervised end-to-end learning of discrete linguistic units for voice conversion,”
2019
Cited alongside, same era.
J. Serrà, S. Pascual, and C. S. Perales, “Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion,” in
2019
Cited alongside, same era.
J. Zhang, Z. Ling, Y. Jiang, L. Liu, C. Liang, and L. Dai, “Improving sequence-to-sequence voice conversion by adding text-supervision,” in
2019
Cited alongside, same era.
J. Zhang, Z. Ling, and L.-R. Dai, “Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations,”
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Later among the works it cites.
2019
Later among the works it cites.
K. Kastner, J. F. Santos, Y. Bengio, and A. Courville, “Representation mixing for tts synthesis,” in
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga
2019
Later among the works it cites.
V. Pratap, A. Hannun, Q. Xu, J. Cai, J. Kahn, G. Synnaeve, V. Liptchinsky, and R. Collobert, “Wav2letter++: A fast open-source speech recognition system,” in
2019
Later among the works it cites.
M. Bińkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, “High fidelity speech synthesis with adversarial networks,” in
2020
Closest in time.
Z.-H. Tan, N. Dehak
2020
Closest in time.