Fetching the paper…
Reading the bibliography…
Using a text description as prompt to guide the generation of text or images (e.g., GPT-3 or DALLE-2) has drawn wide attention recently.
“Language models are few-shot learners,”
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Dhariwal, et al., · 1901
Earlier work this paper cites.
“An objective measure for estimating mos of synthesized speech,”
Chu Min and Peng Hu, · 2001
Earlier work this paper cites.
“Speech quality assessment,”
Philipos C Loizou, · 2011
Earlier work this paper cites.
“Generative adversarial text to image synthesis,”
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, et al., · 2016
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, et al., · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“Token-level ensemble distillation for grapheme-to-phoneme conversion,”
Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, et al., · 2019
Earlier work this paper cites.
“Libritts: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, et al., · 2019
Cited alongside, same era.
“Generative pretraining from pixels,”
Mark Chen, Alec Radford, Rewon Child, Jeff Wu, Heewoo Jun, Prafulla Dhariwal, et al., · 2020
Cited alongside, same era.
“Speaking speed control of end-to-end speech synthesis using sentence-level conditioning,”
Jae-Sung Bae, Hanbin Bae, Young-Sun Joo, Junmo Lee, Gyeong-Hoon Lee, et al., · 2020
Cited alongside, same era.
“Simbert: Integrating retrieval and generation into bert,”
Jianlin Su, · 2020
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Pangu- α \alpha : Large-scale autoregressive pretrained chinese language models with auto-parallel computation,”
“A survey on neural speech synthesis,”
Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu, · 2021
Later among the works it cites.
“Fastpitchformant: Source-filter based decomposed modeling for speech synthesis,”
Taejun Bak, Jae-Sung Bae, Hanbin Bae, Young-Ik Kim, and Hoon-Young Cho, · 2021
Later among the works it cites.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, et al., · 2021
Later among the works it cites.
“Hierarchical text-conditional image generation with clip latents,”
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen, · 2022
Closest in time.
“Naturalspeech: End-to-end text to speech synthesis with human-level quality,”
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, et al., · 2022
Closest in time.
“Unsupervised word-level prosody tagging for controllable speech synthesis,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, et al., · 2021
Cited alongside, same era.
“Cpm-2: Large-scale cost-effective pre-trained language models,”
Zhengyan Zhang, Yuxian Gu, Xu Han, Shengqi Chen, Chaojun Xiao, et al., · 2021
Cited alongside, same era.
Yiwei Guo, Chenpeng Du, and Kai Yu, · 2022
Closest in time.
“P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks,”
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, et al., · 2022
Closest in time.