Fetching the paper…
Reading the bibliography…
The tokenization of speech with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context.
A method for the construction of minimum-redundancy codes
David A Huffman · 1952
Earlier work this paper cites.
Perceptual coding of digital audio
T. Painter and A. Spanias · 2000
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra · 2001
Earlier work this paper cites.
Elements of information theory
MTCAJ Thomas and A Thomas Joy · 2006
Earlier work this paper cites.
Spectral audio signal processing
Julius O Smith · 2011
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
High-quality, low-delay music coding in the Opus codec
Jean-Marc Valin, Gregory Maxwell, Timothy B Terriberry, and Koen Vos · 2016
Earlier work this paper cites.
A non-intrusive short-time objective intelligibility measure
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, and Jesper Jensen · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Chris Donahue, Julian McAuley, and Miller Puckette · 2018
Earlier work this paper cites.
Parallel wavenet: Fast high-fidelity speech synthesis
Aaron Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et al · 2018
Earlier work this paper cites.
webMUSHRA—a comprehensive framework for web-based listening tests
Michael Schoeffler, Sarah Bartoschek, Fabian-Robert Stöter, Marlene Roess, Susanne Westphal, Bernd Edler, and Jürgen Herre · 2018
Earlier work this paper cites.
SDR–half-baked or well done?
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey · 2019
Earlier work this paper cites.
MOSnet: Deep learning based objective assessment for voice conversion
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang · 2019
Earlier work this paper cites.
Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Earlier work this paper cites.
MLS: A large-scale multilingual dataset for speech research
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert · 2020
Earlier work this paper cites.
auraloss: Audio focused loss functions in PyTorch
Christian J. Steinmetz and Joshua D. Reiss · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Conditional sound generation using neural discrete time-frequency representation learning
Xubo Liu, Turab Iqbal, Jinzheng Zhao, Qiushi Huang, Mark D Plumbley, and Wenwu Wang · 2021
Cited alongside, same era.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Vampnet: Music generation via masked acoustic token modeling
Hugo Flores Garcia, Prem Seetharaman, Rithesh Kumar, and Bryan Pardo · 2023
Later among the works it cites.
High-fidelity audio compression with improved RVQGAN
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · 2023
Later among the works it cites.
AudioLDM: Text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Later among the works it cites.
Finite scalar quantization: VQ-VAE made simple
Fabian Mentzer, David C. Minnen, Eirikur Agustsson, and Michael Tschannen · 2023
Later among the works it cites.
Full-band general audio synthesis with score-based diffusion
Santiago Pascual, Gautam Bhattacharya, Chunghsin Yeh, Jordi Pons, and Joan Serrà · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Simon Rouard and Gaëtan Hadjeres · 2021
Cited alongside, same era.
Going deeper with image transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Cited alongside, same era.
MaskGIT: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman · 2022
Cited alongside, same era.
WavLM: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, and Zhengyang et. al. Chen · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Cited alongside, same era.
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Later among the works it cites.
ConvNeXt V2: Co-designing and scaling convnets with masked autoencoders
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie · 2023
Later among the works it cites.
Neural speech coding for real-time communications using constant bitrate scalar quantization
Andreas Brendel, Nicola Pia, Kishan Gupta, Lyonel Behringer, Guillaume Fuchs, and Markus Multrus · 2024
Closest in time.
Moshi: a speech-text foundation model for real-time dialogue
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Single-codec: Single-codebook speech codec towards high-performance speech generation
Hanzhao Li, Liumeng Xue, Haohan Guo, Xinfa Zhu, Yuanjun Lv, Lei Xie, Yunlin Chen, Hao Yin, and Zhifei Li · 2024
Closest in time.
How should we extract discrete audio tokens from self-supervised models?
Pooneh Mousavi, Jarod Duret, Salah Zaiem, Luca Della Libera, Artem Ploujnikov, Cem Subakan, and Mirco Ravanelli · 2024
Closest in time.
StemGen: A music generation model that listens
Julian Parker, Janne Spijkervet, Katerina Kosta, Furkan Yesiler, Boris Kuznetsov, Ju-Chiang Wang, Matt Avent, Jitong Chen, and Duc Le · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
SimpleSpeech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models
Dongchao Yang, Dingdong Wang, Haohan Guo, Xueyuan Chen, Xixin Wu, and Helen Meng · 2024
Closest in time.
Retrieval-augmented text-to-audio generation
Yi Yuan, Haohe Liu, Xubo Liu, Qiushi Huang, Mark D Plumbley, and Wenwu Wang · 2024
Closest in time.