Fetching the paper…
Reading the bibliography…
In speech synthesis and speech enhancement systems, melspectrograms need to be precise in acoustic representations.
Z. Wang, A.C. Bovik, H.R. Sheikh, E.P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE Transactions on Image Processing,
2004
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A shorttime objective intelligibility measure for time-frequency weighted noisy speech,” in IEEE ICASSP
2010
Earlier work this paper cites.
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio, “Generative adversarial nets,” in Advances in neural information processing systems,
2014
Earlier work this paper cites.
Kingma, Diederik P and Ba, Jimmy Lei, “Adam: A method for stochastic optimization,” arXiv preprint
2014
Earlier work this paper cites.
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner,A. W. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” InSSW
2016
Earlier work this paper cites.
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros, “Image-to-image translation with conditional adversarial networks,” arXiv preprint
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton,Y. Wu,R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen,S. Bengio, et al., “Tacotron: A fully end-to-endtext-to-speech synthesis model,” arXiv preprint: 1703.10135
2017
Earlier work this paper cites.
J. Sotelo, S. Mehri, K. Kumar, J. F. San-tos, K. Kastner, A. Courville and Y. Bengio, “Char2wav: End-to-end speech synthesis,” ICLR 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, et al., “Natural tts synthesis by con-ditioning wavenet on mel spectrogram predictions,” arXivpreprint
2017
Cited alongside, same era.
S. Pascual, A. Bonafonte, and J. Serra, “SEGAN: Speech Enhancement Generative Adversarial Network,” in INTERSPEECH
2017
Cited alongside, same era.
T. Kaneko, S. Takaki, H. Kameoka and J. Yamagishi, “Generative adversarial network-based postf
2017
Cited alongside, same era.
D. Michelsanti and Z.-H. Tan, “Conditional generative adversarial networks for speech enhancement and noise-robust speaker verification,” in INTERSPEECH
2017
Cited alongside, same era.
Zhengxin Zhang, Qingjie Liu, and Yunhong Wang, “Road Extraction by Deep Residual U-Net,” arXiv preprint
2017
Cited alongside, same era.
SHAH, Neil, PATIL, Hemant A., et SONI, Meet H. “Time-Frequency Mask-based Speech Enhancement using Convolutional Generative Adversarial Network,” In 2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
2018
Later among the works it cites.
Saito, Y., Takamichi, S., Saruwatari, H. (2018). “Text-to-Speech Synthesis Using STFT Spectra Based on Low-/Multi-Resolution Generative Adversarial Networks,” In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP),
2018
Later among the works it cites.
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, Bryan Catanzaro, “High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs”, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR),
2018
Later among the works it cites.
Leyuan Sheng, Evgeniy N. Pavlovskiy, “Reducing over-smoothness in speech synthesis using Generative Adversarial Networks,” arXiv preprint
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keith Ito, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/
2017
Cited alongside, same era.
W. Ping, K. Peng and J. Chen, “Clarinet: Parallel wave generation in end-to-endtext-to-speech,” arXiv preprint
2018
Cited alongside, same era.
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” arXiv preprint
2018
Cited alongside, same era.
Later among the works it cites.
O. Ernst, Shlomo E. Chazan, S. Gannot and J. Goldberger, “Speech Dereverberation Using Fully Convolutional Networks,” arXiv preprint
2018
Later among the works it cites.
Y. Gao, R. Singh, B. Raj, “Voice impersonation using generative adversarial networks,” IEEE Transactions on Acoustics, Speech and Signal Processing
2018
Later among the works it cites.
Spratley, S.; Beck, D.; Cohn, T, “A Unified Neural Architecture for Instrumental Audio Tasks,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2019
Closest in time.