Fetching the paper…
Reading the bibliography…
While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming.
“Notes on the history of correlation,”
K. Pearson, · 1920
Earlier work this paper cites.
“Algorithm as 136: A k-means clustering algorithm,”
J. A Hartigan and M. A Wong, · 1979
Earlier work this paper cites.
“The proof and measurement of association between two things,”
C. Spearman, · 1987
Earlier work this paper cites.
“An adaptive algorithm for mel-cepstral analysis of speech.,”
T. Fukada, K. Tokuda, T. Kobayashi, et al., · 1992
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,”
A. W Rix, J. G Beerends, M. P Hollier, et al., · 2001
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation,”
K. Papineni, S. Roukos, T. Ward, et al., · 2002
Earlier work this paper cites.
“The blizzard challenge 2008,”
S. King, R. AJ Clark, C. Mayo, et al., · 2008
Earlier work this paper cites.
“The blizzard challenge 2009,”
S. King and V. Karaiskosb, · 2009
Earlier work this paper cites.
“A pitch tracking corpus with evaluation on multipitch tracking scenario,”
Gregor Pirker, Michael Wohlmayr, Stefan Petrik, and Franz Pernkopf, · 2011
Earlier work this paper cites.
“The blizzard challenge 2013–indian language task,”
K. Prahallad, A. Vadapalli, N. Elluru, et al., · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“XSEDE: Accelerating scientific discovery,”
J. Towns, T. Cockerill, M. Dahan, et al., · 2014
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Bridges: a uniquely flexible hpc resource for new communities and data analytics,”
Nicholas A Nystrom, Michael J Levine, Ralph Z Roskies, and J Ray Scott, · 2015
Earlier work this paper cites.
“AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech,”
B. Patton, Y. Agiomyrgiannakis, M. Terry, et al., · 2016
Cited alongside, same era.
“The voice conversion challenge 2016.,”
T. Toda, L.-H. Chen, D. Saito, et al., · 2016
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, et al., · 2017
Cited alongside, same era.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Cited alongside, same era.
“The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,”
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, et al., · 2018
Cited alongside, same era.
“MOSNet: Deep learning based objective assessment for voice conversion,”
“Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,”
C. K A Reddy, V. Gopal, and R. Cutler, · 2021
Later among the works it cites.
“On generative spoken language modeling from raw audio,”
K. Lakhotia, E. Kharitonov, W.-N. Hsu, et al., · 2021
Later among the works it cites.
“Text-free prosody-aware generative spoken language modeling,”
E. Kharitonov, A. Lee, A. Polyak, et al., · 2021
Later among the works it cites.
“Bartscore: Evaluating generated text as text generation,”
W. Yuan, G. Neubig, and P. Liu, · 2021
Later among the works it cites.
“Towards unsupervised learning of speech features in the wild,”
M. Rivière and E. Dupoux, · 2021
Later among the works it cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C.-C. Lo, S.-W. Fu, W.-C. Huang, et al., · 2019
Cited alongside, same era.
“MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance,”
W. Zhao, M. Peyrard, F. Liu, et al., · 2019
Cited alongside, same era.
“Bertscore: Evaluating text generation with bert,”
T. Zhang, V. Kishore, F. Wu, et al., · 2019
Cited alongside, same era.
“Discretalk: Text-to-speech as a machine translation problem,”
T. Hayashi and S. Watanabe, · 2020
Cited alongside, same era.
“The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,”
C. KA Reddy, V. Gopal, R. Cutler, et al., · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, et al., · 2020
Cited alongside, same era.
“Voice conversion challenge 2020intra-lingual semi-parallel and cross-lingual voice conversion,”
Z. Yi, W.-C. Huang, X. Tian, et al., · 2020
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. Tsai, et al., · 2021
Later among the works it cites.
“Superb: Speech processing universal performance benchmark,”
Shu-wen et al. Yang, · 2021
Later among the works it cites.
“SQuId: Measuring speech naturalness in many languages,”
T. Sellam, A. Bapna, J. Camp, et al., · 2022
Closest in time.
“Generalization ability of mos prediction networks,”
E. Cooper, W.-C. Huang, T. Toda, et al., · 2022
Closest in time.
“The voicemos challenge 2022,”
W.-C. Huang, E. Cooper, Y. Tsao, et al., · 2022
Closest in time.
“Audiolm: a language modeling approach to audio generation,”
Z. Borsos, R. Marinier, D. Vincent, et al., · 2022
Closest in time.
“UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022,”
T. Saeki, D. Xin, W. Nakata, et al., · 2022
Closest in time.
“Speech Quality Assessment through MOS using Non-Matching References,”
P. Manocha and A. Kumar, · 2022
Closest in time.
“Back to the Future: Extending the Blizzard Challenge 2013,”
S. L. Maguer, S. King, and N. Harte, · 2022
Closest in time.