Fetching the paper…
Reading the bibliography…
Recent advancements in deep learning led to human-level performance in single-speaker speech synthesis.
B. Sisman, M. Zhang, and H. Li, “A voice conversion framework with tandem feature sparse representation and speaker-adapted wavenet vocoder.” in
1982
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,”
1984
Earlier work this paper cites.
R. McAulay and T. Quatieri, “Speech analysis/synthesis based on a sinusoidal representation,”
1986
Earlier work this paper cites.
L. J. Liu, Z. H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai, “WaveNet vocoder with limited training data for voice conversion.” in
1987
Earlier work this paper cites.
E. Moulines and F. Charpentier, “Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,”
1990
Earlier work this paper cites.
T. Dutoit,
1997
Earlier work this paper cites.
Y. Stylianou, O. Cappé, and E. Moulines, “Continuous probabilistic transform for voice conversion,”
1998
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. De Cheveigne, “Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds,”
1999
Earlier work this paper cites.
J. Kominek and A. W. Black, “The CMU arctic speech databases,” in
2004
Earlier work this paper cites.
P. Taylor,
2009
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,”
2009
Earlier work this paper cites.
S. King, “An introduction to statistical parametric speech synthesis,”
2011
Earlier work this paper cites.
Y. Qian, F. K. Soong, and Z.-J. Yan, “A unified trajectory tiling approach to high quality speech rendering,”
2012
Earlier work this paper cites.
N. Perraudin, P. Balazs, and P. L. Søndergaard, “A fast Griffin-Lim algorithm,” in
2013
Cited alongside, same era.
T. Merritt, R. A. Clark, Z. Wu, J. Yamagishi, and S. King, “Deep neural network-guided unit selection synthesis,” in
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Cited alongside, same era.
V. Christophe, Y. Junichi, and M. Kirsten, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,”
2016
Cited alongside, same era.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder.” in
2017
Cited alongside, same era.
S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in
2018
Later among the works it cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu
2018
Later among the works it cites.
D. Paul, Y. Pantazis, and Y. Stylianou, “Non-parallel voice conversion using weighted generative adversarial networks.” in
2019
Later among the works it cites.
W. Ping, K. Peng, and J. Chen, “ClariNet: Parallel wave generation in end-to-end text-to-speech,” in
2019
Later among the works it cites.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A flow-based generative network for speech synthesis,” in
2019
Later among the works it cites.
J. M. Valin and J. Skoglund, “LPCnet: Improving neural speech synthesis through linear prediction,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” in
2017
Cited alongside, same era.
T. Hayashi, A. Tamamori, K. Kobayashi, K. Takeda, and T. Toda, “An investigation of multi-speaker training for wavenet vocoder,” in
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Cited alongside, same era.
A. Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. Driessche, E. Lockhart, L. Cobo, F. Stimberg
2018
Cited alongside, same era.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in
2018
Cited alongside, same era.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in
2018
Cited alongside, same era.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie
2018
Cited alongside, same era.
2019
Later among the works it cites.
S. Vasquez and M. Lewis, “MelNet: A generative model for audio in the frequency domain,”
2019
Later among the works it cites.
J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, R. Barra-Chicote, A. Moinet, and V. Aggarwal, “Towards achieving robust universal neural vocoding,” in
2019
Later among the works it cites.
J. Park, K. Zhao, K. Peng, and W. Ping, “Multi-speaker end-to-end speech synthesis,”
2019
Later among the works it cites.
Q. Hu, E. Marchi, D. Winarsky, Y. Stylianou, D. Naik, and S. Kajarekar, “Neural text-to-speech adaptation from low quality public recordings,” in
2019
Later among the works it cites.
M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, “Cross-lingual, multi-speaker text-to-speech synthesis using neural speaker embedding,” in
2019
Later among the works it cites.
E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,” in
2020
Closest in time.