Fetching the paper…
Reading the bibliography…
Modern speech synthesis systems have improved significantly, with synthetic speech being indistinguishable from real speech.
“On information and sufficiency”
Solomon Kullback and Richard Leibler · 1951
Earlier work this paper cites.
“GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio”
Guoguo Chen et al · 1965
Earlier work this paper cites.
“Divergence measures based on the Shannon entropy”
J. Lin · 1991
Earlier work this paper cites.
“A metric for distributions with applications to image databases”
Y. Rubner, C. Tomasi and L.J. Guibas · 1998
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks”
Alex Graves, Santiago Fernández, Faustino Gomez and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
“Diagnostic evaluation of synthetic speech using speech recognition”
Miloš Cerňak, Milan Rusko and Marian Trnka · 2009
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books”
Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
“Towards the next generation of web-based experiments: A case study assessing basic audio quality following the ITU-R recommendation BS. 1534 (MUSHRA)”
Michael Schoeffler, Fabian-Robert Stöter, Bernd Edler and Jürgen Herre · 2015
Earlier work this paper cites.
“Hybrid CTC/Attention Architecture for End-to-End Speech Recognition”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John. Hershey and Tomoki Hayashi · 2017
Earlier work this paper cites.
“ESPnet: End-to-End Speech Processing Toolkit”
Shinji Watanabe et al · 2018
Earlier work this paper cites.
“Empirical approaches to measuring the intelligibility of different varieties of English in predicting listener comprehension”
Okim Kang, Ron Thomson and Meghan Moran · 2018
Cited alongside, same era.
“High Fidelity Speech Synthesis with Adversarial Networks”
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis Cobo and Karen Simonyan · 2019
Cited alongside, same era.
“Mosnet: Deep learning based objective assessment for voice conversion”
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao and Hsin-Min Wang · 2019
Cited alongside, same era.
“LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron Weiss, Ye Jia, Zhifeng Chen and Yonghui Wu · 2019
Cited alongside, same era.
“CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92)”, 2019
Junichi Yamagishi, Christophe Veaux and Kirsten MacDonald · 2019
“Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone”
Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Junior, Eren Gölge and Moacir Ponti · 2022
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision”
Alec Radford, Jong Kim, Tao Xu, Greg Brockman, Christine McLeavey and Ilya Sutskever · 2022
Later among the works it cites.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama and Alex Ray · 2022
Later among the works it cites.
“Speechlmscore: Evaluating Speech Generation Using Speech Language Model”
Soumi Maiti, Yifan Peng, Takaaki Saeki and Shinji Watanabe · 2023
Closest in time.
“SQuId: Measuring speech naturalness in many languages”
Thibault Sellam, Ankur Bapna, Joshua Camp, Diana Mackinnon, Ankur Parikh and Jason Riesa · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“On Generative Spoken Language Modeling from Raw Audio”
Kushal Lakhotia et al · 2021
Cited alongside, same era.
“Generalization ability of MOS prediction networks”
Erica Cooper, Wen-Chin Huang, Tomoki Toda and Junichi Yamagishi · 2022
Cited alongside, same era.
“Evaluating and reducing the distance between synthetic and real speech distributions”
Christoph Minixhofer, Ondřej Klejch and Peter Bell · 2022
Cited alongside, same era.
“Styletts: A style-based generative model for natural and diverse text-to-speech synthesis”
Yinghao Li, Cong Han and Nima Mesgarani · 2022
Cited alongside, same era.
“Voicebox: Text-guided multilingual universal speech generation at scale”
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi and Jay Mahadeokar · 2023
Closest in time.
“Neural codec language models are zero-shot text to speech synthesizers”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang and Jinyu Li · 2023
Closest in time.
“A vector quantized approach for text to speech synthesis on real-world spontaneous speech”
Li-Wei Chen, Shinji Watanabe and Alexander Rudnicky · 2023
Closest in time.
“Transformer Based Grapheme-to-Phoneme Conversion”
Sevinj Yolchuyeva, Géza Németh and Bálint Gyires-Tóth · 2099
Closest in time.