Fetching the paper…
Reading the bibliography…
Music codecs are a vital aspect of audio codec research, and ultra low-bitrate compression holds significant importance for music transmission and generation.
“A high-quality speech and audio codec with less than 10-ms delay,”
Jean-Marc Valin, Timothy B Terriberry, Christopher Montgomery, and Gregory Maxwell, · 2009
Earlier work this paper cites.
“Constant-q transform toolbox for music processing,”
Christian Schörkhuber and Anssi Klapuri, · 2010
Earlier work this paper cites.
“Overview of the evs codec architecture,”
Martin Dietz, Markus Multrus, Vaclav Eksler, Vladimir Malenovsky, Erik Norvell, Harald Pobloth, Lei Miao, Zhe Wang, Lasse Laaksonen, Adriana Vasilache, et al., · 2015
Earlier work this paper cites.
“Towards the next generation of web-based experiments: A case study assessing basic audio quality following the itu-r recommendation bs. 1534 (mushra),”
Michael Schoeffler, Fabian-Robert Stöter, Bernd Edler, and Jürgen Herre, · 2015
Earlier work this paper cites.
“Visqol: an objective speech quality model,”
Andrew Hines, Jan Skoglund, Anil C Kokaram, and Naomi Harte, · 2015
Earlier work this paper cites.
“High-quality, low-delay music coding in the opus codec,”
Jean-Marc Valin, Gregory Maxwell, Timothy B Terriberry, and Koen Vos, · 2016
Earlier work this paper cites.
“Subjective evaluation of music compressed with the acer codec compared to aac, mp3, and uncompressed pcm,”
Stuart Cunningham and Iain McGregor, · 2019
Earlier work this paper cites.
“Audio codec enhancement with generative adversarial networks,”
Arijit Biswas and Dai Jia, · 2020
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Earlier work this paper cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy, · 2020
Earlier work this paper cites.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2021
Earlier work this paper cites.
“Mp3: A unified model to map, perceive, predict and plan,”
Sergio Casas, Abbas Sadat, and Raquel Urtasun, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“Postgan: A gan-based post-processor to enhance the quality of coded speech,”
Srikanth Korse, Nicola Pia, Kishan Gupta, and Guillaume Fuchs, · 2022
Cited alongside, same era.
“Flow matching for generative modeling,”
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le, · 2022
Cited alongside, same era.
Siqi Zheng, Luyao Cheng, Yafeng Chen, Hui Wang, and Qian Chen, · 2023
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2023
Later among the works it cites.
“Simple and controllable music generation,”
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez, · 2024
Closest in time.
“An end-to-end approach for chord-conditioned song generation,” 2024
Shuochen Gao, Shun Lei, Fan Zhuo, Hangyu Liu, Feng Liu, Boshi Tang, Qiaochu Huang, Shiyin Kang, and Zhiyong Wu, · 2024
Closest in time.
“Songcreator: Lyrics-based universal song generation,” 2024
Shun Lei, Yixuan Zhou, Boshi Tang, Max W. Y. Lam, Feng Liu, Hangyu Liu, Jingcheng Wu, Shiyin Kang, Zhiyong Wu, and Helen Meng, · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Self-supervised learning with random-projection quantizer for speech recognition,”
Chung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu, and Yonghui Wu, · 2022
Cited alongside, same era.
“Classifier-free diffusion guidance,”
Jonathan Ho and Tim Salimans, · 2022
Cited alongside, same era.
“Audiodec: An open-source streaming high-fidelity neural audio codec,”
Yi-Chiao Wu, Israel D Gebru, Dejan Marković, and Alexander Richard, · 2023
Cited alongside, same era.
“Audiodec: An open-source streaming high-fidelity neural audio codec,”
Yi-Chiao Wu, Israel D Gebru, Dejan Marković, and Alexander Richard, · 2023
Cited alongside, same era.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Cited alongside, same era.
“Scalable diffusion models with transformers,”
William Peebles and Saining Xie, · 2023
Cited alongside, same era.
“Mert: Acoustic music understanding model with large-scale self-supervised training,”
Yizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin, Anton Ragni, Emmanouil Benetos, et al., · 2023
Cited alongside, same era.
“Speechx: Neural codec language model as a versatile speech transformer,”
Xiaofei Wang, Manthan Thakker, Zhuo Chen, Naoyuki Kanda, Sefik Emre Eskimez, Sanyuan Chen, Min Tang, Shujie Liu, Jinyu Li, and Takuya Yoshioka, · 2024
Closest in time.
“Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec,”
Zhihao Du, Shiliang Zhang, Kai Hu, and Siqi Zheng, · 2024
Closest in time.
“Apcodec: A neural audio codec with parallel amplitude and phase spectrum encoding and decoding,”
Yang Ai, Xiao-Hang Jiang, Ye-Xin Lu, Hui-Peng Du, and Zhen-Hua Ling, · 2024
Closest in time.
“High-fidelity audio compression with improved rvqgan,”
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2024
Closest in time.
“Semanticodec: An ultra low bitrate semantic audio codec for general sound,”
Haohe Liu, Xuenan Xu, Yi Yuan, Mengyue Wu, Wenwu Wang, and Mark D Plumbley, · 2024
Closest in time.
“Seed-tts: A family of high-quality versatile speech generation models,”
Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, et al., · 2024
Closest in time.
“Audioldm 2: Learning holistic audio generation with self-supervised pretraining,”
Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D Plumbley, · 2024
Closest in time.
“High-fidelity audio compression with improved rvqgan,”
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2024
Closest in time.