Fetching the paper…
Reading the bibliography…
We introduce a novel sequence-to-sequence (seq2seq) voice conversion (VC) model based on the Transformer architecture with text-to-speech (TTS) pretraining.
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai, “WaveNet Vocoder with Limited Training Data for Voice Conversion,” in
1987
Earlier work this paper cites.
Y. Stylianou, O. Cappe, and E. Moulines, “Continuous probabilistic transform for voice conversion,”
1998
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. de Cheveigné, “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based F0 extraction: Possible role of a repetitive structure in sounds,”
1999
Earlier work this paper cites.
J. Kominek and A. W. Black, “The CMU ARCTIC speech databases,” in
2004
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice Conversion Based on Maximum-Likelihood Estimation of Spectral Parameter Trajectory,”
2007
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to Sequence Learning with Neural Networks,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: An ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications,”
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
H. Miyoshi, Y. Saito, S. Takamichi, and H. Saruwatari, “Voice Conversion Using Sequence-to-Sequence Learning of Context Posterior Probabilities,” in
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is All you Need,” in
2017
Earlier work this paper cites.
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin, “Convolutional Sequence to Sequence Learning,” in
2017
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in
2017
Cited alongside, same era.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent wavenet vocoder,” in
2017
Cited alongside, same era.
T. Hayashi, A. Tamamori, K. Kobayashi, K. Takeda, and T. Toda, “An investigation of multi-speaker training for WaveNet vocoder,” in
2017
Cited alongside, same era.
2018
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition,” in
2019
Closest in time.
G. Zhao, S. Ding, and R. Gutierrez-Osuna, “Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams,” in
2019
Closest in time.
2019
Closest in time.
F. Biadsy, R. J. Weiss, P. J. Moreno, D. Kanvesky, and Y. Jia, “Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech Separation,” in
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
L. Cross Vila, C. Escolano, J. A. R. Fonollosa, and M. R. Costa-Jussà, “End-to-End Speech Translation with the Transformer,” in
2018
Cited alongside, same era.
H. Tachibana, K. Uenoyama, and S. Aihara, “Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention,” in
2018
Cited alongside, same era.
S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in
2018
Cited alongside, same era.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. Lopez Moreno, and Y. Wu, “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS Synthesis by Conditioning WaveNet on MEL Spectrogram Predictions,” in
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-End Speech Processing Toolkit,” in
2018
Cited alongside, same era.
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, “ATTS2S-VC: Sequence-to-sequence Voice Conversion with Attention and Context Preservation Mechanisms,” in
2019
Cited alongside, same era.
2019
Closest in time.
M. A. D. Gangi, M. Negri, R. Cattoni, R. Dessi, and M. Turchi, “Enhancing Transformer for End-to-end Speech-to-Text Translation,” in
2019
Closest in time.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural Speech Synthesis with Transformer Network,” in
2019
Closest in time.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie, C. Gulcehre, A. van den Oord, O. Vinyals, and N. de Freitas, “Sample efficient adaptive text-to-speech,” in
2019
Closest in time.
2019
Closest in time.
H.-T. Luong and J. Yamagishi, “Bootstrapping non-parallel voice conversion from speaker-adaptive text-to-speech,” in
2019
Closest in time.
M. Zhang, X. Wang, F. Fang, H. Li, and J. Yamagishi, “Joint Training Framework for Text-to-Speech and Voice Conversion Using Multi-Source Tacotron and WaveNet,” in
2019
Closest in time.
Munich Artificial Intelligence Laboratories GmbH, “The M-AILABS speech dataset,” 2019, accessed 30 November 2019. [Online]. Available:
2019
Closest in time.