Fetching the paper…
Reading the bibliography…
Historically, most speech models in machine-learning have used the mel-spectrogram as a speech representation.
B.-H. Juang and A. Gray, “Multiple stage vector quantization for speech coding,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 1982, p. 597–600
1982
Earlier work this paper cites.
M. Bisani and H. Ney, “Bootstrap estimates for confidence intervals in ASR performance evaluation,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , vol. 1, 2004, pp. I–409
2004
Earlier work this paper cites.
A. Hines, J. Skoglund, A. C. Kokaram, and N. Harte, “ViSQOL: an objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, no. 1, pp. 1–18, 2015
2015
Earlier work this paper cites.
J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 11, pp. 2009–2022, 2016
2016
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2018, pp. 4779–4783
2018
Earlier work this paper cites.
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR - half-baked or well done?” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2019
2019
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. Conf. on Neural Information Process. Systems (NeurIPS) , 2020
2020
Earlier work this paper cites.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Proc. Interspeech . ISCA, Oct. 2020
2020
Earlier work this paper cites.
A. Łańcucki, “FastPitch: Parallel text-to-speech with pitch prediction,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2021, pp. 6588–6592
2021
Earlier work this paper cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and high-quality end-to-end text to speech,” in Proc. International Conference on Learning Representations (ICLR) , 2021
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
R. Badlani, A. Łancucki, K. J. Shih, R. Valle, W. Ping, and B. Catanzaro, “One TTS alignment to rule them all,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved RVQGAN,” in Proc. Conf. on Neural Information Process. Systems (NeurIPS) , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Kumar, K. Tan, Z. Ni, P. Manocha, X. Zhang, E. Henderson, and B. Xu, “TorchAudio-Squim: Reference-less speech quality and intelligibility measures in torchaudio,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
H. Wei, R. Xie, H. Cheng, L. Feng, B. An, and Y. Li, “Mitigating neural network overconfidence with logit normalization,” in Proc. International Conference on Machine Learning (ICML) , 2022
2022
Cited alongside, same era.
Y. Ren, X. Tan, T. Qin, Z. Zhao, and T.-Y. Liu, “Revisiting over-smoothness in text to speech,” in Proc. 60th Annual Meeting of the Association for Computational Linguistics , 2022
2022
Cited alongside, same era.
S. gil Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, “BigVGAN: A universal neural vocoder with large-scale training,” in Proc. International Conference on Learning Representations (ICLR) , 2023
2023
Cited alongside, same era.
Y.-C. Wu, I. D. Gebru, D. Marković, and A. Richard, “AudioDec: An open-source streaming high-fidelity neural audio codec,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2023
2023
Cited alongside, same era.
E. Harper, S. Majumdar, O. Kuchaiev, L. Jason, Y. Zhang, E. Bakhturina, V. Noroozi, S. Subramanian, K. Nithin, H. Jocelyn, F. Jia, J. Balam, X. Yang, M. Livne, Y. Dong, S. Naren, and B. Ginsburg, “Nemo: a toolkit for conversational ai and large language models.” [Online]. Available: https://github.com/NVIDIA/NeMo
Cited in the paper.
P. Neekhara, S. Hussain, S. Ghosh, J. Li, R. Valle, R. Badlani, and B. Ginsburg, “Improving robustness of LLM-based speech synthesis by learning monotonic alignment,” Proc. Interspeech , 2024, to appear
2024
Closest in time.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
F. Mentzer, D. Minnen, E. Agustsson, and M. Tschannen, “Finite scalar quantization: VQ-VAE made simple,” in Proc. International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
K. C. Puvvada, N. R. Koluguri, K. Dhawan, J. Balam, and B. Ginsburg, “Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2024
2024
Closest in time.
NVIDIA, “STT En Fast Conformer-Transducer XLarge,” https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/stt_en_fastconformer_transducer_xlarge
2024
Closest in time.