Fetching the paper…
Reading the bibliography…
Neural audio codecs have recently gained popularity because they can represent audio signals with high fidelity at very low bitrates, making it feasible to use language modeling approaches for audio generation and understanding.
Multiple stage vector quantization for speech coding
Biing-Hwang Juang and A Gray · 1982
Earlier work this paper cites.
Vector quantization
Robert Gray · 1984
Earlier work this paper cites.
Understanding pitch perception as a hierarchical process with top-down modulation
Emili Balaguer-Ballester, Nicholas R Clark, Martin Coath, Katrin Krumbholz, and Susan L Denham · 2009
Earlier work this paper cites.
Hierarchical processing for speech in human auditory cortex and beyond
Jonathan E Peelle, Ingrid Johnsrude, and Matthew H Davis · 2010
Earlier work this paper cites.
Definition of the Opus Audio Codec
Jean-Marc Valin, Koen Vos, and Timothy B. Terriberry · 2012
Earlier work this paper cites.
Method for the subjective assessment of intermediate quality level of audio systems
B Series · 2014
Earlier work this paper cites.
Can we automatically transform speech recorded on common consumer devices in real-world environments into professional production quality speech?—a dataset, insights, and challenges
Gautham J. Mysore · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
End-to-end optimized speech coding with deep neural networks
Srihari Kankanahalli · 2018
Cited alongside, same era.
A hierarchical latent vector model for learning long-term structure in music
Adam Roberts, Jesse Engel, Colin Raffel, Curtis Hawthorne, and Douglas Eck · 2018
Cited alongside, same era.
The challenge of realistic music generation: modelling raw audio at scale
Sander Dieleman, Aaron Van Den Oord, and Karen Simonyan · 2018
Cited alongside, same era.
webmushra-a comprehensive framework for web-based listening tests
Michael Schoeffler, Sarah Bartoschek, Fabian-Robert Stoter, Marlene Roess, Susanne Westphal, Bernd Edler, and Jurgen Herre · 2018
Cited alongside, same era.
Low bit-rate speech coding with vq-vae and a wavenet decoder
Cristina Gârbacea, Aäron van den Oord, Yazhe Li, Felicia SC Lim, Alejandro Luebs, Oriol Vinyals, and Thomas C Walters · 2019
Cited alongside, same era.
Generating diverse high-fidelity images with vq-vae-2
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Later among the works it cites.
Univnet: A neural vocoder with multi-resolution spectrogram discriminators for high-fidelity waveform generation
Won Jang, Dan Lim, Jaesam Yoon, Bongwan Kim, and Juntae Kim · 2021
Later among the works it cites.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Later among the works it cites.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals · 2019
Cited alongside, same era.
Musdb18-hq - an uncompressed version of musdb18, 2019
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner · 2019
Cited alongside, same era.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Cited alongside, same era.
Bigvgan: A universal neural vocoder with large-scale training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon · 2023
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · 2024
Closest in time.
Speechtokenizer: Unified speech tokenizer for speech language models
Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu · 2024
Closest in time.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez · 2024
Closest in time.