Fetching the paper…
Reading the bibliography…
Mellotron is a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data.
“Midi: musical instrument digital interface,”
Robert A Moog, · 1986
Earlier work this paper cites.
“Musicxml for notation and analysis,”
Michael Good, · 2001
Earlier work this paper cites.
“Yin, a fundamental frequency estimator for speech and music,”
Alain De Cheveigné and Hideki Kawahara, · 2002
Earlier work this paper cites.
“A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,”
Tomohiro Nakatani, Shigeaki Amano, Toshio Irino, Kentaro Ishizuka, and Tadahisa Kondo, · 2008
Earlier work this paper cites.
“Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,”
Wei Chu and Abeer Alwan, · 2009
Earlier work this paper cites.
“Melody extraction from polyphonic music signals using pitch contour characteristics,”
Justin Salamon and Emilia Gómez, · 2012
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Singing voice synthesis based on deep neural networks,”
Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda, · 2016
Cited alongside, same era.
“Deep voice 2: Multi-speaker neural text-to-speech,”
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou, · 2017
Cited alongside, same era.
“Deep voice 3: Scaling text-to-speech with convolutional sequence learning,”
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2017
Cited alongside, same era.
“Montreal forced aligner: Trainable text-speech alignment using kaldi.,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Cited alongside, same era.
“The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito, · 2017
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Later among the works it cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Later among the works it cites.
“Adversarially trained end-to-end korean singing voice synthesis system,”
Juheon Lee, Hyeong-Seok Choi, Chang-Bin Jeon, Junghyun Koo, and Kyogu Lee, · 2019
Closest in time.
“Jasper: An end-to-end convolutional neural acoustic model,”
Jason Li, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen, Huyen Nguyen, and Ravi Teja Gadde, · 2019
Closest in time.
“Libritts: A corpus derived from librispeech for text-to-speech,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J Weiss, Rob Clark, and Rif A Saurous, · 2018
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A Saurous, · 2018
Cited alongside, same era.
“Gentle forced aligner,”
R. M. Ochshorn and M. Hawkins,
Cited in the paper.
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Closest in time.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Closest in time.