Fetching the paper…
Reading the bibliography…
We present a meta-learning approach for adaptive text-to-speech (TTS) with few data.
The formation of learning sets
H. F. Harlow · 1949
Earlier work this paper cites.
An Introduction to Text-to-speech Synthesis
T. Dutoit · 1997
Earlier work this paper cites.
Yin, a fundamental frequency estimator for speech and music
A. De Cheveigné and H. Kawahara · 2002
Earlier work this paper cites.
Text-to-Speech Synthesis
P. Taylor · 2009
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Learning to learn
S. Thrun and L. Pratt · 2012
Earlier work this paper cites.
Fast speaker adaptation of hybrid nn/hmm model for speech recognition based on discriminative learning of speaker code
O. Abdel-Hamid and H. Jiang · 2013
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, A. Rabinovich, et al · 2015
Earlier work this paper cites.
DRAW: A recurrent neural network for image generation
K. Gregor, I. Danihelka, A. Graves, D. Rezende, and D. Wierstra · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
The kestrel tts text normalization system
P. Ebden and R. Sproat · 2015
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Fast, compact, and high quality LSTM-RNN based statistical parametric speech synthesizers for mobile devices
H. Zen, Y. Agiomyrgiannakis, N. Egberts, F. Henderson, and P. Szczepaniak · 2016
Earlier work this paper cites.
Meta-learning with memory-augmented neural networks
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap · 2016
Earlier work this paper cites.
Matching networks for one shot learning
O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al · 2016
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, B. Shillingford, and N. De Freitas · 2016
Cited alongside, same era.
Optimization as a model for few-shot learning
S. Ravi and H. Larochelle · 2016
Cited alongside, same era.
One-shot generalization in deep generative models
D. J. Rezende, S. Mohamed, I. Danihelka, K. Gregor, and D. Wierstra · 2016
Cited alongside, same era.
Pixel recurrent neural networks
A. Van Oord, N. Kalchbrenner, and K. Kavukcuoglu · 2016
Cited alongside, same era.
World: a vocoder-based high-quality speech synthesis system for real-time applications
M. Morise, F. Yokomori, and K. Ozawa · 2016
Cited alongside, same era.
Deep voice 2: Multi-speaker neural text-to-speech
A. Gibiansky, S. Arik, G. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou · 2017
Later among the works it cites.
Deep voice: Real-time neural text-to-speech
S. Ö. Arık, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, et al · 2017
Later among the works it cites.
Char2wav: End-to-end speech synthesis
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio · 2017
Later among the works it cites.
CSTR VCTK corpus: English multi-speaker corpus for CSTR Voice Cloning Toolkit, 2017
C. Veaux, J. Yamagishi, K. MacDonald, et al · 2017
Later among the works it cites.
Generalized end-to-end loss for speaker verification
L. Wan, Q. Wang, A. Papir, and I. L. Moreno · 2018
Closest in time.
Observe and look further: Achieving consistent performance on atari
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Redefining the linguistic context feature set for hmm and dnn tts through position and parsing
R. Dall, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda · 2016
Cited alongside, same era.
Parallel WaveNet: Fast high-fidelity speech synthesis
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. v. d. Driessche, E. Lockhart, L. C. Cobo, F. Stimberg, et al · 2017
Cited alongside, same era.
Deep speaker feature learning for text-independent speaker verification
L. Li, Y. Chen, Y. Shi, Z. Tang, and D. Wang · 2017
Cited alongside, same era.
Attentive recurrent comparators
P. Shyam, S. Gupta, and A. Dukkipati · 2017
Cited alongside, same era.
Learning to learn without gradient descent by gradient descent
Y. Chen, M. W. Hoffman, S. G. Colmenarejo, M. Denil, T. P. Lillicrap, M. Botvinick, and N. Freitas · 2017
Cited alongside, same era.
Fast adaptation in generative models with generative matching networks
S. Bartunov and D. P. Vetrov · 2017
Cited alongside, same era.
T. Pohlen, B. Piot, T. Hester, M. G. Azar, D. Horgan, D. Budden, G. Barth-Maron, H. van Hasselt, J. Quan, M. Večerík, et al · 2018
Closest in time.
Playing hard exploration games by watching youtube
Y. Aytar, T. Pfaff, D. Budden, T. L. Paine, Z. Wang, and N. de Freitas · 2018
Closest in time.
One-shot imitation from observing humans via domain-adaptive meta-learning
T. Yu, C. Finn, A. Xie, S. Dasari, T. Zhang, P. Abbeel, and S. Levine · 2018
Closest in time.
Few-shot autoregressive density estimation: Towards learning to learn distributions
S. Reed, Y. Chen, T. Paine, A. van den Oord, S. M. Eslami, D. Rezende, O. Vinyals, and N. de Freitas · 2018
Closest in time.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous · 2018
Closest in time.
Deep voice 3: 2000-speaker neural text-to-speech
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller · 2018
Closest in time.
Voiceloop: Voice fitting and synthesis via a phonological loop
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani · 2018
Closest in time.
Fitting new speakers based on a short untranscribed sample
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf · 2018
Closest in time.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno, et al · 2018
Closest in time.
Neural voice cloning with a few samples
S. O. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou · 2018
Closest in time.