Fetching the paper…
Reading the bibliography…
While audio quality is a key performance metric for various audio processing tasks, including generative modeling, its objective measurement remains a challenge.
Notes on the history of correlation
Pearson, K · 1920
Earlier work this paper cites.
Limitations of perceptual evaluation of speech quality on voip systems
Manjunath, T · 2009
Earlier work this paper cites.
Robustness of speech quality metrics to background noise and network degradations: Comparing visqol, pesq and polqa
Hines, A., Skoglund, J., Kokaram, A., and Harte, N · 2013
Earlier work this paper cites.
Visqol: an objective speech quality model
Hines, A., Skoglund, J., Kokaram, A. C., and Harte, N · 2015
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D. P. W., Gemmeke, J. F., Jansen, A., Moore, C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., Slaney, M., Weiss, R., and Wilson, K · 2017
Earlier work this paper cites.
Tackling toxic online communication with recurrent capsule networks
Deshmukh, S. and Rade, R · 2018
Earlier work this paper cites.
Fr \ \backslash ’echet audio distance: A metric for evaluating music enhancement algorithms
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2018
Earlier work this paper cites.
Training supervised speech separation system to improve stoi and pesq directly
Zhang, H., Zhang, X., and Gao, G · 2018
Earlier work this paper cites.
Non-intrusive speech quality assessment using neural networks
Avila, A. R., Gamper, H., Reddy, C., Cutler, R., Tashev, I., and Gehrke, J · 2019
Earlier work this paper cites.
Learning with learned loss function: Speech enhancement with quality-net to improve perceptual evaluation of speech quality
Fu, S.-W., Liao, C.-F., and Tsao, Y · 2019
Earlier work this paper cites.
AudioCaps: Generating Captions for Audios in The Wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Mosnet: Deep learning-based objective assessment for voice conversion
Lo, C.-C., Fu, S.-W., Huang, W.-C., Wang, X., Yamagishi, J., Tsao, Y., and Wang, H.-M · 2019
Earlier work this paper cites.
The blizzard challenge 2019
Wu, Z., Xie, Z., and King, S · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech
Zen, H., Dang, V., Clark, R., Zhang, Y., Weiss, R. J., Jia, Y., Chen, Z., and Wu, Y · 2019
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., et al · 2020
Cited alongside, same era.
A differentiable perceptual audio metric learned from just noticeable differences
Manocha, P., Finkelstein, A., Zhang, R., Bryan, N. J., Mysore, G. J., and Jin, Z · 2020
Cited alongside, same era.
Cdpam: Contrastive learning for perceptual audio similarity
Manocha, P., Jin, Z., Zhang, R., and Finkelstein, A · 2021
Cited alongside, same era.
NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets
Mittag, G., Naderi, B., Chehadi, A., and Möller, S · 2021
Cited alongside, same era.
Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors
Reddy, C. K., Gopal, V., and Cutler, R · 2021
Cited alongside, same era.
Styletts: A style-based generative model for natural and diverse text-to-speech synthesis
Li, Y. A., Han, C., and Mesgarani, N · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Progressive distillation for fast sampling of diffusion models
Salimans, T. and Ho, J · 2022
Later among the works it cites.
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Icassp 2021 deep noise suppression challenge
Reddy, C. K. A., Dubey, H., Gopal, V., Cutler, R., Braun, S., Gamper, H., Aichner, R., and Srinivasan, S · 2021
Cited alongside, same era.
Sesqa: semi-supervised learning for speech quality assessment
Serrà, J., Pons, J., and Pascual, S · 2021
Cited alongside, same era.
Objective measures of perceptual audio quality reviewed: An evaluation of their application domain dependence
Torcoli, M., Kastner, T., and Herre, J · 2021
Cited alongside, same era.
Narle: Natural language models using reinforcement learning with emotion feedback
Zhou, R., Deshmukh, S., Greer, J., and Lee, C · 2021
Cited alongside, same era.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone
Casanova, E., Weber, J., Shulby, C. D., Junior, A. C., Gölge, E., and Ponti, M. A · 2022
Cited alongside, same era.
Generalization ability of mos prediction networks
Cooper, E., Huang, W.-C., Toda, T., and Yamagishi, J · 2022
Cited alongside, same era.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Cited alongside, same era.
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al · 2023
Later among the works it cites.
Evaluating speech synthesis by training recognizers on synthetic speech
Alharthi, D., Sharma, R., Dhamyal, H., Maiti, S., Raj, B., and Singh, R · 2023
Later among the works it cites.
Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies
Chen*, K., Wu*, Y., Liu*, H., Nezhurina, M., Berg-Kirkpatrick, T., and Dubnov, S · 2023
Later among the works it cites.
A vector quantized approach for text to speech synthesis on real-world spontaneous speech
Chen, L.-W., Watanabe, S., and Rudnicky, A · 2023
Later among the works it cites.
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A · 2023
Later among the works it cites.
Clap learning audio concepts from natural language supervision
Elizalde, B., Deshmukh, S., Ismail, M. A., and Wang, H · 2023
Later among the works it cites.
Compa: Addressing the gap in compositional reasoning in audio-language models
Ghosh, S., Seth, A., Kumar, S., Tyagi, U., Evuru, C. K., Ramaneswaran, S., Sakshi, S., Nieto, O., Duraiswami, R., and Manocha, D · 2023
Later among the works it cites.
Adapting frechet audio distance for generative music evaluation
Gui, A., Gamper, H., Braun, S., and Emmanouilidou, D · 2023
Later among the works it cites.
Voicebox: Text-guided multilingual universal speech generation at scale
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., et al · 2023
Later among the works it cites.
Speechlmscore: Evaluating speech generation using speech language model
Maiti, S., Peng, Y., Saeki, T., and Watanabe, S · 2023
Later among the works it cites.
Mubert, 2023
Mubert-Inc · 2023
Later among the works it cites.