Fetching the paper…
Reading the bibliography…
Automatic methods to predict listener opinions of synthesized speech remain elusive since listeners, systems being evaluated, characteristics of the speech, and even the instructions given and the rating scale all vary from test to test.
“SWITCHBOARD: telephone speech corpus for research and development,”
John J Godfrey, Edward C Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“XIMERA: A new TTS from ATR based on corpus-based technologies,”
Hisashi Kawai, Tomoki Toda, Jinfu Ni, Minoru Tsuzaki, and Keiichi Tokuda, · 2004
Earlier work this paper cites.
“The Fisher Corpus: a resource for the next generations of speech-to-text,”
Christopher Cieri, David Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“The Blizzard Challenge 2008,”
Vasilis Karaiskos, Simon King, Robert AJ Clark, and Catherine Mayo, · 2008
Earlier work this paper cites.
“The Blizzard Challenge 2009,”
Alan W Black, Simon King, and Keiichi Tokuda, · 2009
Earlier work this paper cites.
“The Blizzard Challenge 2010,”
Simon King and Vasilis Karaiskos, · 2010
Earlier work this paper cites.
“The Blizzard Challenge 2011,”
Simon King and Vasilis Karaiskos, · 2011
Earlier work this paper cites.
“BABEL: IARPA solicitation IARPA-BAA-11-02,”
Mary Harper, · 2011
Earlier work this paper cites.
“The Blizzard Challenge 2013,”
Simon King and Vasilis Karaiskos, · 2013
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“The Blizzard Challenge 2016,”
Simon King and Vasilis Karaiskos, · 2016
Earlier work this paper cites.
“The Voice Conversion Challenge 2016,”
Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, and Junichi Yamagishi, · 2016
Cited alongside, same era.
“Analysis of the Voice Conversion Challenge 2016 evaluation results.,”
Mirjam Wester, Zhizheng Wu, and Junichi Yamagishi, · 2016
Cited alongside, same era.
“The Voice Conversion Challenge 2018: Promoting development of parallel and nonparallel methods,”
Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, and Zhenhua Ling, · 2018
Cited alongside, same era.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai, · 2018
Cited alongside, same era.
“A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,”
Xin Wang, Jaime Lorenzo-Trueba, Shinji Takaki, Lauri Juvela, and Junichi Yamagishi, · 2018
Cited alongside, same era.
“Predictions of subjective ratings and spoofing assessments of Voice Conversion Challenge 2020 submissions,”
Rohan Kumar Das, Tomi Kinnunen, Wen-Chin Huang, Zhenhua Ling, Junichi Yamagishi, Yi Zhao, Xiaohai Tian, and Tomoki Toda, · 2020
Later among the works it cites.
“ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,” 2020
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan, · 2020
Later among the works it cites.
“ASVspoof 2019: a large-scale public database of synthesized, converted and replayed speech,”
Xin Wang, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas Evans, Md Sahidullah, Ville Vestman, Tomi Kinnunen, Kong Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao, Hsin-Min Wang, Sébastien Le Maguer, Markus Becker, Fergus Henderson, Rob Clark, Yu Zhang, Quan Wang, Ye Jia, Kai Onuma, Koji Mushika, Takashi Kaneda, Yuan Jiang, Li-Juan Liu, Yi-Chiao Wu, Wen-Chin Huang, Tomoki Toda, Kou Tanaka, Hirokazu Kameoka, Ingmar Steiner, Driss Matrouf, Jean-François Bonastre, Avashna Govender, Srikanth Ronanki, Jing-Xuan Zhang, and Zhen-Hua Ling, · 2020
Later among the works it cites.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“MOSNet: deep learning-based objective assessment for voice conversion,”
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang, · 2019
Cited alongside, same era.
“ASVspoof 2019: future horizons in spoofed and fake audio detection,”
Massimiliano Todisco, Xin Wang, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Tomi H Kinnunen, and Kong Aik Lee, · 2019
Cited alongside, same era.
“The blizzard challenge 2019,”
Zhizheng Wu, Zhihang Xie, and Simon King, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Comparison of speech representations for automatic quality estimation in multi-speaker text-to-speech synthesis,”
Jennifer Williams, Joanna Rownicka, Pilar Oplustil, and Simon King, · 2020
Cited alongside, same era.
“Voice Conversion Challenge 2020 — intra-lingual semi-parallel and cross-lingual voice conversion —,”
Zhao Yi, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhenhua Ling, and Tomoki Toda, · 2020
Cited alongside, same era.
Later among the works it cites.
“Common Voice: a massively-multilingual speech corpus,”
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber, · 2020
Later among the works it cites.
“MLS: a large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert, · 2020
Later among the works it cites.
“How do voices from past speech synthesis challenges compare today?,”
Erica Cooper and Junichi Yamagishi, · 2021
Closest in time.
“HuBERT: self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.
“MBNET: MOS prediction for synthesized speech with mean-bias network,”
Yichong Leng, Xu Tan, Sheng Zhao, Frank Soong, Xiang-Yang Li, and Tao Qin, · 2021
Closest in time.
“Utilizing self-supervised representations for MOS prediction,”
Wei-Cheng Tseng, Chien-yu Huang, Wei-Tsung Kao, Yist Y Lin, and Hung-yi Lee, · 2021
Closest in time.
“SUPERB: speech processing universal performance benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al., · 2021
Closest in time.