Fetching the paper…
Reading the bibliography…
Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling.
Principles of digital audio
Ken C Pohlmann, · 2000
Earlier work this paper cites.
“Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,”
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, · 2001
Earlier work this paper cites.
“A short-time objective intelligibility measure for time-frequency weighted noisy speech,”
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen, · 2010
Earlier work this paper cites.
“Crema-d: Crowd-sourced emotional multimodal actors dataset,”
Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma, · 2014
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Quesst2014: Evaluating query-by-example speech search in a zero-resource setting with real-life queries,”
Xavier Anguera, Luis-J Rodriguez-Fuentes, Andi Buzo, Florian Metze, Igor Szöke, and Mikel Penagarikano, · 2015
Earlier work this paper cites.
“ESC: Dataset for Environmental Sound Classification,”
Karol J. Piczak, · 2015
Earlier work this paper cites.
“Voxceleb: A large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Earlier work this paper cites.
“The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R Livingstone and Frank A Russo, · 2018
Earlier work this paper cites.
“Speech model pre-training for end-to-end spoken language understanding,”
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio, · 2019
Earlier work this paper cites.
“Librimix: An open-source dataset for generalizable speech separation,”
Joris Cosentino, Manuel Pariente, Samuele Cornell, Antoine Deleforge, and Emmanuel Vincent, · 2020
Earlier work this paper cites.
“Gunshots recorded in an open field using ipod touch devices,”
Seth Cooper and Steven Shaw, · 2020
Earlier work this paper cites.
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck, · 2020
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2021
Earlier work this paper cites.
“VoxLingua107: a dataset for spoken language recognition,”
Jörgen Valk and Tanel Alumäe, · 2021
Earlier work this paper cites.
“Semi-supervised spoken language understanding via self-supervised speech and language model pretraining,”
Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, and James R. Glass, · 2021
Earlier work this paper cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu et al., · 2021
Earlier work this paper cites.
“Audiogen: Textually guided audio generation,”
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi, · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“FSD50K: an open dataset of human-labeled sound events,”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra, · 2022
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2022
Cited alongside, same era.
“ludlows/python-pesq: supporting for multiprocessing features,” May 2022
Miao Wang, Christoph Boeddeker, Rafael G. Dantas, and ananda seelan, · 2022
Cited alongside, same era.
“Clap learning audio concepts from natural language supervision,”
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang, · 2023
Later among the works it cites.
“Towards audio language modeling-an overview,”
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, Kai-wei Chang, Ho-Lam Chung, Alexander H Liu, and Hung-yi Lee, · 2024
Closest in time.
“Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,”
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al., · 2024
Closest in time.
Shengpeng Ji, Minghui Fang, Ziyue Jiang, Rongjie Huang, Jialung Zuo, Shulei Wang, and Zhou Zhao, · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Masked autoencoders that listen,”
Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer, · 2022
Cited alongside, same era.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al., · 2023
Cited alongside, same era.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Cited alongside, same era.
“Musiclm: Generating music from text,”
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al., · 2023
Cited alongside, same era.
“Soundstorm: Efficient parallel audio generation,”
Zalán Borsos, Matt Sharifi, Damien Vincent, Eugene Kharitonov, Neil Zeghidour, and Marco Tagliasacchi, · 2023
Cited alongside, same era.
“Audiodec: An open-source streaming high-fidelity neural audio codec,”
Yi-Chiao Wu, Israel D Gebru, Dejan Marković, and Alexander Richard, · 2023
Cited alongside, same era.
“Hifi-codec: Group-residual vector quantization for high fidelity audio codec,”
Dongchao Yang, Songxiang Liu, Rongjie Huang, Jinchuan Tian, Chao Weng, and Yuexian Zou, · 2023
Cited alongside, same era.
Yi-Chiao Wu, Dejan Marković, Steven Krenn, Israel D Gebru, and Alexander Richard, · 2024
Closest in time.
“Srcodec: Split-residual vector quantization for neural speech codec,”
Youqiang Zheng, Weiping Tu, Li Xiao, and Xinmeng Xu, · 2024
Closest in time.
“Supercodec: A neural speech codec with selective back-projection network,”
Youqiang Zheng, Weiping Tu, Li Xiao, and Xinmeng Xu, · 2024
Closest in time.
“Lightcodec: A high fidelity neural audio codec with low computation complexity,”
Liang Xu, Jing Wang, Jianqian Zhang, and Xiang Xie, · 2024
Closest in time.
“Generative de-quantization for neural speech codec via latent diffusion,”
Haici Yang, Inseon Jang, and Minje Kim, · 2024
Closest in time.
“Semanticodec: An ultra low bitrate semantic audio codec for general sound,”
Haohe Liu, Xuenan Xu, Yi Yuan, Mengyue Wu, Wenwu Wang, and Mark D Plumbley, · 2024
Closest in time.
“Apcodec: A neural audio codec with parallel amplitude and phase spectrum encoding and decoding,”
Yang Ai, Xiao-Hang Jiang, Ye-Xin Lu, Hui-Peng Du, and Zhen-Hua Ling, · 2024
Closest in time.
“Single-codec: Single-codebook speech codec towards high-performance speech generation,”
Hanzhao Li, Liumeng Xue, Haohan Guo, Xinfa Zhu, Yuanjun Lv, Lei Xie, Yunlin Chen, Hao Yin, and Zhifei Li, · 2024
Closest in time.
Haohan Guo, Fenglong Xie, Dongchao Yang, Hui Lu, Xixin Wu, and Helen Meng, · 2024
Closest in time.
“Uniaudio 1.5: Large language model-driven audio codec is a few-shot audio task learner,”
Dongchao Yang, Haohan Guo, Yuanyuan Wang, Rongjie Huang, Xiang Li, Xu Tan, Xixin Wu, and Helen Meng, · 2024
Closest in time.
“The interspeech 2024 challenge on speech processing using discrete units,”
Xuankai Chang, Jiatong Shi, Jinchuan Tian, Yuning Wu, Yuxun Tang, Yihan Wu, Shinji Watanabe, Yossi Adi, Xie Chen, and Qin Jin, · 2024
Closest in time.
“Dasb–discrete audio and speech benchmark,”
Pooneh Mousavi, Luca Della Libera, Jarod Duret, Artem Ploujnikov, Cem Subakan, and Mirco Ravanelli, · 2024
Closest in time.
“Codec-superb: An in-depth analysis of sound codec models,”
Haibin Wu, Ho-Lam Chung, Yi-Cheng Lin, Yuan-Kuei Wu, Xuanjun Chen, Yu-Chi Pai, Hsiu-Hsuan Wang, Kai-Wei Chang, Alexander H Liu, and Hung-yi Lee, · 2024
Closest in time.