Fetching the paper…
Reading the bibliography…
In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals.
Single-ended speech quality measurement using machine learning methods
Tiago H Falk and W-Y Chan. 2006 · 1947
Earlier work this paper cites.
A study of complexity and quality of speech waveform coders
JM Tribolet, Peter Noll, B McDermott, and R Crochiere. 1978 · 1978
Earlier work this paper cites.
Prediction of perceived phonetic distance from critical-band spectra: A first step
Dennis Klatt. 1982 · 1982
Earlier work this paper cites.
Objective measures for speech quality testing
Thomas P Barnwell III, MA Clements, and SR Quackenbush. 1988 · 1988
Earlier work this paper cites.
Conversation analysis
Charles Goodwin and John Heritage. 1990 · 1990
Earlier work this paper cites.
Mel-cepstral distance measure for objective speech quality assessment
Robert Kubichek. 1993 · 1993
Earlier work this paper cites.
Output-based objective speech quality
Jin Liang and Robert Kubichek. 1994 · 1994
Earlier work this paper cites.
Telephone transmission quality subjective opinion tests. a method for subjective performance assessment of the quality of speech voice output devices
ITU-T Recommendation. 1994 · 1994
Earlier work this paper cites.
5-khz-bandwidth speech coder at 4-8 kbit/s
Noboru Harada and Hitoshi Ohmuro. 1999 · 1999
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra. 2001 · 2001
Earlier work this paper cites.
Spoken dialogue technology: enabling the conversational user interface
Michael F McTear. 2002 · 2002
Earlier work this paper cites.
Bss_eval toolbox user guide–revision 2.0
Cédric Févotte, Rémi Gribonval, and Emmanuel Vincent. 2005 · 2005
Earlier work this paper cites.
Coherence and the speech intelligibility index
James M Kates and Kathryn H Arehart. 2005 · 2005
Earlier work this paper cites.
Evaluation of objective quality measures for speech enhancement
Yi Hu and Philipos C Loizou. 2007 · 2007
Earlier work this paper cites.
Analysis of a simplified normalized covariance measure based on binary weighting functions for predicting the intelligibility of noise-suppressed speech
Fei Chen and Philipos C Loizou. 2010 · 2010
Earlier work this paper cites.
A non-intrusive quality and intelligibility measure of reverberant and dereverberated speech
Tiago H Falk, Chenxi Zheng, and Wai-Yip Chan. 2010 · 2010
Earlier work this paper cites.
Shulei Ji, Jing Luo, and Xinyu Yang. 2020 · 2011
Earlier work this paper cites.
The Kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. 2011 · 2011
Earlier work this paper cites.
An algorithm for intelligibility prediction of time–frequency weighted noisy speech
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen. 2011 · 2011
Earlier work this paper cites.
Speech Enhancement: Theory and Practice , 2nd edition
Philipos C. Loizou. 2013 · 2013
Earlier work this paper cites.
Mir_eval: A transparent implementation of common mir metrics
Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel PW Ellis, and C Colin Raffel. 2014 · 2014
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
The voice conversion challenge 2016
Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, and Junichi Yamagishi. 2016 · 2016
Earlier work this paper cites.
Noisy speech database for training speech enhancement algorithms and tts models, 2016 [sound]
Cassia Valentini-Botinhao. 2017 · 2016
Earlier work this paper cites.
FMA: A dataset for music analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, and Xavier Bresson. 2017 · 2017
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson. 2017 · 2017
Earlier work this paper cites.
Bias and statistical significance in evaluating speech synthesis with mean opinion scores
Andrew Rosenberg and Bhuvana Ramabhadran. 2017 · 2017
Earlier work this paper cites.
Spectral feature mapping with mimic loss for robust speech recognition
Deblin Bagchi, Peter Plantinga, Adam Stiff, and Eric Fosler-Lussier. 2018 · 2018
Earlier work this paper cites.
Spoken conversational AI in video games: Emotional dialogue management increases user engagement
Jamie Fraser, Ioannis Papaioannou, and Oliver Lemon. 2018 · 2018
Earlier work this paper cites.
TaSNet: Time-domain audio separation network for real-time, single-channel speech separation
Yi Luo and Nima Mesgarani. 2018 · 2018
Earlier work this paper cites.
ESPnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai. 2018 · 2018
Earlier work this paper cites.
Multi-speaker emotional acoustic modeling for CNN-based speech synthesis
Heejin Choi, Sangjun Park, Jinuk Park, and Minsoo Hahn. 2019 · 2019
Earlier work this paper cites.
Present in body or just in mind: differences in social presence and emotion regulation in live vs. virtual singing experiences
Daisy Fancourt and Andrew Steptoe. 2019 · 2019
Earlier work this paper cites.
An efficient model for estimating subjective quality of separated audio source signals
Thorsten Kastner and Jürgen Herre. 2019 · 2019
Earlier work this paper cites.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. 2019 · 2019
Earlier work this paper cites.
Conversational AI: An overview of methodologies, applications & future scope
Pradnya Kulkarni, Ameya Mahabaleshwarkar, Mrunalini Kulkarni, Nachiket Sirsikar, and Kunal Gadgil. 2019 · 2019
Earlier work this paper cites.
MOSNet: Deep learning-based objective assessment for voice conversion
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang. 2019 · 2019
Earlier work this paper cites.
Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation
Yi Luo and Nima Mesgarani. 2019 · 2019
Cited alongside, same era.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
Dominik Roblek, Kevin Kilgour, Matt Sharifi, and Mauricio Zuluaga. 2019 · 2019
Cited alongside, same era.
Visqol v3: An open source production ready objective speech and audio metric
Michael Chinen, Felicia SC Lim, Jan Skoglund, Nikita Gureev, Feargus O’Gorman, and Andrew Hines. 2020 · 2020
Cited alongside, same era.
Predictions of subjective ratings and spoofing assessments of voice conversion challenge 2020 submissions
Rohan Kumar Das, Tomi Kinnunen, Wen-Chin Huang, Zhen-Hua Ling, Junichi Yamagishi, Zhao Yi, Xiaohai Tian, and Tomoki Toda. 2020 · 2020
Cited alongside, same era.
ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan. 2020 · 2020
Aquatk: An audio quality assessment toolkit
Ashvala Vinay and Alexander Lerch. 2023 · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui*, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023 · 2023
Later among the works it cites.
Visinger2: High-fidelity end-to-end singing voice synthesis enhanced by digital signal processing synthesizer
Yongmao Zhang, Heyang Xue, Hanzhao Li, Lei Xie, Tingwei Guo, Ruixiong Zhang, and Caixia Gong. 2023b · 2023
Later among the works it cites.
Melotts: High-quality multi-lingual multi-accent text-to-speech
Wenliang Zhao, Xumin Yu, and Zengyi Qin. 2023 · 2023
Later among the works it cites.
Chattts: A generative speech model for daily dialogue
2Noise. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Xiaoicesing: A high-quality and integrated singing voice synthesis system
Peiling Lu, Jie Wu, Jian Luan, Xu Tan, and Li Zhou. 2020 · 2020
Cited alongside, same era.
Reliable fidelity and diversity metrics for generative models
Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, and Jaejun Yoo. 2020 · 2020
Cited alongside, same era.
Convolutive transfer function invariant sdr training criteria for multi-channel reverberant speech separation
Christoph Boeddeker, Wangyou Zhang, Tomohiro Nakatani, Keisuke Kinoshita, Tsubasa Ochiai, Marc Delcroix, Naoyuki Kamo, Yanmin Qian, and Reinhold Haeb-Umbach. 2021 · 2021
Cited alongside, same era.
Metricgan+: An improved version of metricgan for speech enhancement
Szu-Wei Fu, Cheng Yu, Tsun-An Hsieh, Peter Plantinga, Mirco Ravanelli, Xugang Lu, and Yu Tsao. 2021 · 2021
Cited alongside, same era.
Espnet2-tts: Extending the edge of tts research
Tomoki Hayashi, Ryuichi Yamamoto, Takenori Yoshimura, Peter Wu, Jiatong Shi, Takaaki Saeki, Yooncheol Ju, Yusuke Yasuda, Shinnosuke Takamichi, and Shinji Watanabe. 2021 · 2021
Cited alongside, same era.
Neural dubber: Dubbing for videos according to scripts
Chenxu Hu, Qiao Tian, Tingle Li, Wang Yuping, Yuxuan Wang, and Hang Zhao. 2021 · 2021
Cited alongside, same era.
Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks
Jee-weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-jin Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, and Nicholas Evans. 2021 · 2021
Cited alongside, same era.
Kaito Baba, Wataru Nakata, Yuki Saito, and Hiroshi Saruwatari. 2024 · 2024
Closest in time.
MusicLDM: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies
Ke Chen, Yusong Wu, Haohe Liu, Marianna Nezhurina, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2024a · 2024
Closest in time.
https://github.com/modelscope/clearervoice-studio
ClearerVoice-Studio. 2023 · 2024
Closest in time.
Whisperspeech: A speech processing toolkit
Collabora. 2024 · 2024
Closest in time.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2024 · 2024
Closest in time.
Ai-based affective music generation systems: A review of methods and challenges
Adyasha Dash and Kathleen Agres. 2024 · 2024
Closest in time.
PAM: Prompting audio-language models for audio quality assessment
Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde, Hannes Gamper, Mahmoud Al Ismail, Rita Singh, Bhiksha Raj, and Huaming Wang. 2024 · 2024
Closest in time.
Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, et al. 2024 · 2024
Closest in time.
Measuring audio prompt adherence with distribution-based embedding distances
Maarten Grachten. 2024 · 2024
Closest in time.
Adapting frechet audio distance for generative music evaluation
Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Emmanouilidou. 2024 · 2024
Closest in time.
The VoiceMOS challenge 2024: Beyond speech quality prediction
Wen-Chin Huang, Szu-Wei Fu, Erica Cooper, Ryandhimas E Zezario, Tomoki Toda, Hsin-Min Wang, Junichi Yamagishi, and Yu Tsao. 2024b · 2024
Closest in time.
Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling
Shengpeng Ji, Ziyue Jiang, Wen Wang, Yifu Chen, Minghui Fang, Jialong Zuo, Qian Yang, Xize Cheng, Zehan Wang, Ruiqi Li, et al. 2024 · 2024
Closest in time.
log-wmse: A weighted log spectral loss for audio quality estimation
Iver Jordal. 2023 · 2024
Closest in time.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. 2024 · 2024
Closest in time.
Parler-tts
Yoach Lacombe, Vaibhav Srivastav, and Sanchit Gandhi. 2024 · 2024
Closest in time.
The limits of the mean opinion score for speech synthesis evaluation
Sébastien Le Maguer, Simon King, and Naomi Harte. 2024 · 2024
Closest in time.
PL-TTS: A generalizable prompt-based diffusion tts augmented by large language model
Shuhua Li, Qirong Mao, and Jiatong Shi. 2024 · 2024
Closest in time.
Audioldm 2: Learning holistic audio generation with self-supervised pretraining
Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley. 2024 · 2024
Closest in time.
emotion2vec: Self-supervised pre-training for speech emotion representation
Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, ShiLiang Zhang, and Xie Chen. 2024 · 2024
Closest in time.
NOMAD: Unsupervised learning of perceptual embeddings for speech enhancement and non-matching reference audio quality assessment
Alessandro Ragano, Jan Skoglund, and Andrew Hines. 2024a · 2024
Closest in time.
Speechbertscore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics
Takaaki Saeki, Soumi Maiti, Shinnosuke Takamichi, Shinji Watanabe, and Hiroshi Saruwatari. 2024 · 2024
Closest in time.
ESPnet-Codec: Comprehensive training and evaluation of neural codecs for audio, music, and speech
Jiatong Shi, Jinchuan Tian, Yihan Wu, Jee-weon Jung, Jia Qi Yip, Yoshiki Masuyama, William Chen, Yuning Wu, Yuxun Tang, Massa Baali, et al. 2024 · 2024
Closest in time.
Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier
SileroTeam. 2024 · 2024
Closest in time.
https://github.com/stability-ai/stable-audio-metrics
Stability-AI. 2023 · 2024
Closest in time.
SingMOS: An extensive open-source singing voice dataset for MOS prediction
Yuxun Tang, Jiatong Shi, Yuning Wu, and Qin Jin. 2024 · 2024
Closest in time.
Espnet-spk: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
Jee weon Jung, Wangyou Zhang, Jiatong Shi, Zakaria Aldeneh, Takuya Higuchi, Alex Gichamba, Barry-John Theobald, Ahmed Hussen Abdelaziz, and Shinji Watanabe. 2024 · 2024
Closest in time.
Toksing: Singing voice synthesis based on discrete tokens
Yuning Wu, Chunlei Zhang, Jiatong Shi, Yuxun Tang, Shan Yang, and Qin Jin. 2024c · 2024
Closest in time.
Bigcodec: Pushing the limits of low-bitrate neural speech codec
Detai Xin, Xu Tan, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2024 · 2024
Closest in time.
Dongchao Yang, Dingdong Wang, Haohan Guo, Xueyuan Chen, Xixin Wu, and Helen Meng. 2024 · 2024
Closest in time.
Emotivoice: a multi-voice and prompt-controlled tts engine
NetEase Youdao. 2024 · 2024
Closest in time.
Visinger2+: End-to-end singing voice synthesis augmented by self-supervised learning representation
Yifeng Yu, Jiatong Shi, Yuning Wu, and Shinji Watanabe. 2024 · 2024
Closest in time.
Meta Audiobox Aesthetics: Unified automatic quality assessment for speech, music, and sound
Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoffman, Brian Ellis, Apoorv Vyas, Bowen Shi, Sanyuan Chen, Matt Le, Nick Zacharov, et al. 2025 · 2025
Closest in time.