Fetching the paper…
Reading the bibliography…
This paper presents FunCodec, a fundamental neural speech codec toolkit, which is an extension of the open-source speech processing toolkit FunASR.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, and et al., · 2011
Earlier work this paper cites.
“Definition of the opus audio codec,”
Jean-Marc Valin, Koen Vos, and Timothy B. Terriberry, · 2012
Earlier work this paper cites.
“Overview of the EVS codec architecture,”
Martin Dietz, Markus Multrus, Vaclav Eksler, and et al., · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Visqol: an objective speech quality model,”
Andrew Hines, Jan Skoglund, Anil C Kokaram, and Naomi Harte, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, and et al., · 2017
Earlier work this paper cites.
“Neural discrete representation learning,”
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu, · 2017
Earlier work this paper cites.
“Automatic differentiation in pytorch,”
Adam Paszke, Sam Gross, Soumith Chintala, and et al., · 2017
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“Aishell-2: Transforming mandarin asr research into industrial scale,”
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu, · 2018
Earlier work this paper cites.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, and et al., · 2019
Cited alongside, same era.
“Libritts: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, and et al., · 2019
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Seanet: A multi-modal speech enhancement network,”
Marco Tagliasacchi, Yunpeng Li, Karolis Misiunas, and Dominik Roblek, · 2020
Cited alongside, same era.
“Real-time speech frequency bandwidth extension,”
Yunpeng Li, Marco Tagliasacchi, Oleg Rybakov, Victor Ungureanu, and Dominik Roblek, · 2021
Cited alongside, same era.
“Low-bitrate redundancy coding of speech using a rate-distortion-optimized variational autoencoder,”
Jean-Marc Valin, Jan Büthe, and Ahmed Mustafa, · 2023
Closest in time.
“LMCodec: A low bitrate speech codec with causal transformer models,”
Teerapat Jenrungrot, Michael Chinen, W. Bastiaan Kleijn, and et al., · 2023
Closest in time.
“Neural feature predictor and discriminative residual coding for low-bitrate speech coding,”
Haici Yang, Wootaek Lim, and Minje Kim, · 2023
Closest in time.
“High quality audio coding with mdctnet,”
Grant Davidson, Mark Vinton, Per Ekstrand, and et al., · 2023
Closest in time.
“End-to-end neural audio coding in the mdct domain,”
Hyungseob Lim, Jihyun Lee, Byeong Hyeon Kim, Inseon Jang, and Hong-Goo Kang, · 2023
Closest in time.
“Neural codec language models are zero-shot text to speech synthesizers,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, and et al., · 2021
Cited alongside, same era.
“Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,”
Guoguo Chen, Shuzhou Chai, Guanbo Wang, and et al., · 2021
Cited alongside, same era.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech,”
Yi Ren, Ming Lei, Zhiying Huang, and et al., · 2022
Cited alongside, same era.
“Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition,”
Binbin Zhang, Hang Lv, Pengcheng Guo, and et al., · 2022
Cited alongside, same era.
Chengyi Wang, Sanyuan Chen, Yu Wu, and et al., · 2023
Closest in time.
“Viola: Unified codec language models for speech recognition, synthesis, and translation,”
Tianrui Wang, Long Zhou, Ziqiang Zhang, and et al., · 2023
Closest in time.
“Audiopalm: A large language model that can speak and listen,”
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, and et al., · 2023
Closest in time.
“AudioDec: An open-source streaming high-fidelity neural audio codec,”
Yi-Chiao Wu, Israel D. Gebru, Dejan Markovic, and Alexander Richard, · 2023
Closest in time.
“High-fidelity audio compression with improved rvqgan,”
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2023
Closest in time.
“Encodec_trainer,”
Michael Meuleman, · 2023
Closest in time.