Fetching the paper…
Reading the bibliography…
Audio signals are sampled at high temporal resolutions, and learning to synthesize audio requires capturing structure across a range of timescales.
A scale for the measurement of the psychological magnitude pitch
Stanley Smith Stevens, John Volkmann, and Edwin B Newman · 1937
Earlier work this paper cites.
Remaking speech
Homer Dudley · 1939
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones
Eric Moulines and Francis Charpentier · 1990
Earlier work this paper cites.
TIMIT acoustic-phonetic continuous speech corpus
John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, David S Pallett, Nancy L Dahlgren, and Victor Zue · 1993
Earlier work this paper cites.
Unit selection in a concatenative speech synthesis system using a large speech database
Andrew J Hunt and Alan W Black · 1996
Earlier work this paper cites.
Simultaneous modeling of phonetic and prosodic parameters, and characteristic conversion for hmm-based text-to-speech systems
Takayoshi Yoshimura · 2002
Earlier work this paper cites.
Speech synthesis based on hidden Markov models
Keiichi Tokuda, Yoshihiko Nankaku, Tomoki Toda, Heiga Zen, Junichi Yamagishi, and Keiichiro Oura · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Deep learning for acoustic modeling in parametric speech generation: A systematic review of existing techniques and future trends
Zhen-Hua Ling, Shi-Yin Kang, Heiga Zen, Andrew Senior, Mike Schuster, Xiao-Jun Qian, Helen M Meng, and Li Deng · 2015
Earlier work this paper cites.
Learning the speech front-end with raw waveform CLDNNs
Tara N Sainath, Ron J Weiss, Andrew Senior, Kevin W Wilson, and Oriol Vinyals · 2015
Earlier work this paper cites.
WORLD: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Earlier work this paper cites.
Deconvolution and checkerboard artifacts
Augustus Odena, Vincent Dumoulin, and Chris Olah · 2016
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2016
Earlier work this paper cites.
Improved techniques for training GANs
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
A note on the evaluation of generative models
Lucas Theis, Aäron van den Oord, and Matthias Bethge · 2016
Cited alongside, same era.
WaveNet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, and Alan W Black · 2016
Cited alongside, same era.
Deep Voice 2: Multi-speaker neural text-to-speech
Least squares generative adversarial networks
Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley · 2017
Later among the works it cites.
SampleRNN: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2017
Later among the works it cites.
Conditional generative adversarial networks for speech enhancement and noise-robust speaker verification
Daniel Michelsanti and Zheng-Hua Tan · 2017
Later among the works it cites.
SEGAN: Speech enhancement generative adversarial network
Santiago Pascual, Antonio Bonafonte, and Joan Serrà · 2017
Later among the works it cites.
Learning from simulated and unsupervised images through adversarial training
Ashish Shrivastava, Tomas Pfister, Oncel Tuzel, Joshua Susskind, Wenda Wang, and Russell Webb · 2017
Later among the works it cites.
Char2Wav: End-to-end speech synthesis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sercan Arik, Gregory Diamos, Andrew Gibiansky, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou · 2017
Cited alongside, same era.
Wasserstein GAN
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Cited alongside, same era.
BEGAN: Boundary equilibrium generative adversarial networks
David Berthelot, Tom Schumm, and Luke Metz · 2017
Cited alongside, same era.
Synthetic speech commands dataset
Johannes Buchner · 2017
Cited alongside, same era.
Deep cross-modal audio-visual generation
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan, and Chenliang Xu · 2017
Cited alongside, same era.
Neural audio synthesis of musical notes with WaveNet autoencoders
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi · 2017
Cited alongside, same era.
SVSGAN: Singing voice separation via generative adversarial network
Zhe-Cheng Fan, Yen-Lin Lai, and Jyh-Shing Roger Jang · 2017
Cited alongside, same era.
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio · 2017
Later among the works it cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Later among the works it cites.
Bird recordings
Peter Boesman · 2018
Closest in time.
Exploring speech enhancement with generative adversarial networks for robust speech recognition
Chris Donahue, Bo Li, and Rohit Prabhavalkar · 2018
Closest in time.
Progressive growing of GANs for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen · 2018
Closest in time.
Deep Voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Closest in time.
Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, et al · 2018
Closest in time.
Speech commands: A dataset for limited-vocabulary speech recognition
Pete Warden · 2018
Closest in time.
GANSynth: Adversarial neural audio synthesis
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts · 2019
Closest in time.