Fetching the paper…
Reading the bibliography…
High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion.
Prolate spheroidal wave functions, fourier analysis and uncertainty—ii
Henry J Landau and Henry O Pollak · 1961
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Unit selection in a concatenative speech synthesis system using a large speech database
Andrew J Hunt and Alan W Black · 1996
Earlier work this paper cites.
An hmm-based system for automatic segmentation and alignment of speech
Kåre Sjölander · 2003
Earlier work this paper cites.
Hidden markov models for grapheme to phoneme conversion
Paul Taylor · 2005
Earlier work this paper cites.
Straight, exploitation of the other aspect of vocoder: Perceptually isomorphic decomposition of speech sounds
Hideki Kawahara · 2006
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2006
Earlier work this paper cites.
Deepsinger: Singing voice synthesis with data mined from the web
Yi Ren, Xu Tan, Tao Qin, Jian Luan, Zhou Zhao, and Tie-Yan Liu · 2007
Earlier work this paper cites.
Comparison of different implementations of mfcc
Fang Zheng, Guoliang Zhang, and Zhanjiang Song · 2009
Earlier work this paper cites.
The heisenberg uncertainty principle and the nyquist-shannon sampling theorem
Pierre A Millette · 2013
Earlier work this paper cites.
Expression control in singing voice synthesis: Features, approaches, evaluation, and challenges
Marti Umbert, Jordi Bonada, Masataka Goto, Tomoyasu Nakano, and Johan Sundberg · 2015
Earlier work this paper cites.
High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network
Lauri Juvela, Bajibabu Bollepalli, Manu Airaksinen, and Paavo Alku · 2016
Earlier work this paper cites.
World: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Cited alongside, same era.
Singing voice synthesis based on deep neural networks
Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Deep voice: Real-time neural text-to-speech
Sercan Ömer Arik, Mike Chrzanowski, Adam Coates, Gregory Frederick Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Y Ng, Jonathan Raiman, et al · 2017
Cited alongside, same era.
A neural parametric singing synthesizer modeling timbre and expression from natural songs
Merlijn Blaauw and Jordi Bonada · 2017
Cited alongside, same era.
Adversarially trained end-to-end korean singing voice synthesis system
Juheon Lee, Hyeong-Seok Choi, Chang-Bin Jeon, Junghyun Koo, and Kyogu Lee · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Singing voice synthesis based on convolutional neural networks
Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda · 2019
Later among the works it cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Later among the works it cites.
Token-level ensemble distillation for grapheme-to-phoneme conversion
Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, Tao Qin, and Tie-Yan Liu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Least squares generative adversarial networks
Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Emphasis: An emotional phoneme-based acoustic model for speech synthesis system
Hao Li, Yongguo Kang, and Zhenyu Wang · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan · 2019
Cited alongside, same era.
Wgansing: A multi-voice singing voice synthesizer based on the wasserstein-gan
Pritish Chandna, Merlijn Blaauw, Jordi Bonada, and Emilia Gómez · 2019
Cited alongside, same era.
Singing voice synthesis based on generative adversarial networks
Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda · 2019
Cited alongside, same era.
Synthesising expressiveness in peking opera via duration informed attention network
Yusong Wu, Shengchen Li, Chengzhu Yu, Heng Lu, Chao Weng, Liqiang Zhang, and Dong Yu · 2019
Later among the works it cites.
Sequence-to-sequence singing synthesis using the feed-forward transformer
Merlijn Blaauw and Jordi Bonada · 2020
Closest in time.
Yu Gu, Xiang Yin, Yonghui Rao, Yuan Wan, Benlai Tang, Yang Zhang, Jitong Chen, Yuxuan Wang, and Zejun Ma · 2020
Closest in time.
Espnet-tts: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan · 2020
Closest in time.
Xiaoicesing: A high-quality and integrated singing voice synthesis system
Peiling Lu, Jie Wu, Jian Luan, Xu Tan, and Li Zhou · 2020
Closest in time.
Fast and high-quality singing voice synthesis system based on convolutional neural networks
Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda · 2020
Closest in time.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Closest in time.