Fetching the paper…
Reading the bibliography…
Neural audio codecs are initially introduced to compress audio data into compact codes to reduce transmission latency.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Rfc 6716: Definition of the opus audio codec,” 2012
Jean-Marc Valin et al., · 2012
Earlier work this paper cites.
“Overview of the evs codec architecture,”
Martin Dietz et al., · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani et al., · 2017
Earlier work this paper cites.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar et al., · 2019
Earlier work this paper cites.
“Seanet: A multi-modal speech enhancement network,”
Marco Tagliasacchi et al., · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Earlier work this paper cites.
“Neural networks fail to learn periodic functions and how to fix it,”
Liu Ziyin, Tilman Hartwig, and Masahito Ueda, · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” 2020
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski et al., · 2020
Earlier work this paper cites.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour et al., · 2021
Earlier work this paper cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu et al., · 2021
Earlier work this paper cites.
“W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,”
Yu-An Chung et al., · 2021
Earlier work this paper cites.
“Speech resynthesis from discrete disentangled self-supervised representations,”
Adam Polyak et al., · 2021
Earlier work this paper cites.
“On generative spoken language modeling from raw audio,”
Kushal Lakhotia et al., · 2021
Earlier work this paper cites.
“Text-free prosody-aware generative spoken language modeling,”
Eugene Kharitonov et al., · 2021
Earlier work this paper cites.
“High fidelity neural audio compression,”
Alexandre Défossez et al., · 2022
Earlier work this paper cites.
“Audiogen: Textually guided audio generation,”
Felix Kreuk et al., · 2022
Earlier work this paper cites.
“Bigvgan: A universal neural vocoder with large-scale training,”
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon, · 2022
Cited alongside, same era.
“Mulan: A joint embedding of music audio and natural language,”
Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, and Daniel PW Ellis, · 2022
Cited alongside, same era.
“Phonetic Analysis of Self-supervised Representations of English Speech,”
Dan Wells, Hao Tang, and Korin Richmond, · 2022
Cited alongside, same era.
“Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation,”
Sravya Popuri et al., · 2022
Cited alongside, same era.
“Unity: Two-pass direct speech-to-speech translation with discrete units,”
Hirofumi Inaguma et al., · 2022
“Uniaudio: An audio foundation model toward universal audio generation,”
Dongchao Yang et al., · 2023
Later among the works it cites.
“Lauragpt: Listen, attend, understand, and regenerate audio with gpt,”
Qian Chen et al., · 2023
Later among the works it cites.
“Speechx: Neural codec language model as a versatile speech transformer,”
Xiaofei Wang et al., · 2023
Later among the works it cites.
“Simple and controllable music generation,”
Jade Copet et al., · 2023
Later among the works it cites.
“Stack-and-delay: a new codebook pattern for music generation,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks,”
Kai-Wei Chang et al., · 2022
Cited alongside, same era.
“Soundstorm: Efficient parallel audio generation,”
Zalán Borsos et al., · 2023
Cited alongside, same era.
“Audiodec: An open-source streaming high-fidelity neural audio codec,”
Yi-Chiao Wu et al., · 2023
Cited alongside, same era.
“Hifi-codec: Group-residual vector quantization for high fidelity audio codec,”
Dongchao Yang et al., · 2023
Cited alongside, same era.
“Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec,”
Zhihao Du, Shiliang Zhang, Kai Hu, and Siqi Zheng, · 2023
Cited alongside, same era.
“Speechtokenizer: Unified speech tokenizer for speech large language models,”
Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu, · 2023
Cited alongside, same era.
“High-fidelity audio compression with improved rvqgan,”
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2023
Cited alongside, same era.
Gael Le Lan et al., · 2023
Later among the works it cites.
Rohan Anil et al., · 2023
Later among the works it cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al., · 2023
Later among the works it cites.
“Megabyte: Predicting million-byte sequences with multiscale transformers,”
Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis, · 2023
Later among the works it cites.
“Generative spoken dialogue language modeling,”
Tu Anh Nguyen et al., · 2023
Later among the works it cites.
“Textually pretrained speech language models,”
Michael Hassid et al., · 2023
Later among the works it cites.
“Seamlessm4t-massively multilingual & multimodal machine translation,”
Loïc Barrault et al., · 2023
Later among the works it cites.
“Seamless: Multilingual expressive and streaming speech translation,”
Loïc Barrault et al., · 2023
Later among the works it cites.
“Speechprompt v2: Prompt tuning for speech classification tasks,”
Kai-Wei Chang et al., · 2023
Later among the works it cites.
“Speechgen: Unlocking the generative power of speech language models with prompts,”
Haibin Wu, Kai-Wei Chang, Yuan-Kuei Wu, and Hung-yi Lee, · 2023
Later among the works it cites.
“An exploration of in-context learning for speech language model,”
Ming-Hao Hsu et al., · 2023
Later among the works it cites.
“Towards general-purpose text-instruction-guided voice conversion,”
Chun-Yi Kuan, Chen-An Li, et al., · 2023
Later among the works it cites.
Chien-yu Huang, Ke-Han Lu, et al., · 2023
Later among the works it cites.