Fetching the paper…
Reading the bibliography…
Audio codec models are widely used in audio communication as a crucial technique for compressing audio into discrete representations.
“Speech analysis and synthesis by linear prediction of the speech wave,”
Bishnu S Atal and Suzanne L Hanauer, · 1971
Earlier work this paper cites.
“Multiple stage vector quantization for speech coding,”
Biing-Hwang Juang and A Gray, · 1982
Earlier work this paper cites.
“Vector quantization,”
Robert Gray, · 1984
Earlier work this paper cites.
“A review of vector quantization techniques,”
A Vasuki and PT Vanathi, · 2006
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron Van Den Oord, Oriol Vinyals, et al., · 2017
Earlier work this paper cites.
“Wavenet based low rate speech coding,”
W Bastiaan Kleijn, Felicia SC Lim, Alejandro Luebs, Jan Skoglund, Florian Stimberg, Quan Wang, and Thomas C Walters, · 2018
Earlier work this paper cites.
“A real-time wideband neural vocoder at 1.6 kb/s using lpcnet,”
Jean-Marc Valin and Jan Skoglund, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Earlier work this paper cites.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2021
Cited alongside, same era.
“Taming visually guided sound generation,”
Vladimir Iashin and Esa Rahtu, · 2021
Cited alongside, same era.
“Fastdiff: A fast conditional diffusion model for high-quality speech synthesis,”
Rongjie Huang, Max WY Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, and Zhou Zhao, · 2022
Cited alongside, same era.
“Prodiff: Progressive fast diffusion model for high-quality text-to-speech,”
Rongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu, Chenye Cui, and Yi Ren, · 2022
Cited alongside, same era.
“Diffsound: Discrete diffusion model for text-to-sound generation,”
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu, · 2022
Cited alongside, same era.
“Data2vec: A general framework for self-supervised learning in speech, vision and language,”
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli, · 2022
Later among the works it cites.
“Disentangling speech from surroundings in a neural audio codec,”
Ahmed Omran, Neil Zeghidour, Zalán Borsos, Félix de Chaumont Quitry, Malcolm Slaney, and Marco Tagliasacchi, · 2022
Later among the works it cites.
“Architecture for variable bitrate neural speech codec with configurable computation complexity,”
Tejas Jayashankar, Thilo Koehler, Kaustubh Kalgaonkar, Zhiping Xiu, Jilong Wu, Ju Lin, Prabhav Agrawal, and Qing He, · 2022
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Closest in time.
“Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi, · 2022
Cited alongside, same era.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour, · 2022
Cited alongside, same era.
“Transpeech: Speech-to-speech translation with bilateral perturbation,”
Rongjie Huang, Zhou Zhao, Jinglin Liu, Huadai Liu, Yi Ren, Lichao Zhang, and Jinzheng He, · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
Dongchao Yang, Songxiang Liu, Rongjie Huang, Guangzhi Lei, Chao Weng, Helen Meng, and Dong Yu, · 2023
Closest in time.
“Musiclm: Generating music from text,”
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al., · 2023
Closest in time.
“Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,”
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao, · 2023
Closest in time.
“A vector quantized approach for text to speech synthesis on real-world spontaneous speech,”
Li-Wei Chen, Shinji Watanabe, and Alexander Rudnicky, · 2023
Closest in time.
“Audiogpt: Understanding and generating speech, music, sound, and talking head,”
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al., · 2023
Closest in time.