Fetching the paper…
Reading the bibliography…
We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks.
A real-time wideband neural vocoder at 1.6 kb/s using lpcnet
Jean-Marc Valin and Jan Skoglund · 1903
Earlier work this paper cites.
Speech analysis and synthesis by linear prediction of the speech wave
Bishnu S Atal and Suzanne L Hanauer · 1971
Earlier work this paper cites.
Source coding algorithms for fast data compression
Richard Clark Pasco · 1976
Earlier work this paper cites.
Universal modeling and coding
Jorma Rissanen and Glen Langdon · 1981
Earlier work this paper cites.
Multiple stage vector quantization for speech coding
Biing-Hwang Juang and A Gray · 1982
Earlier work this paper cites.
Vector quantization
Robert Gray · 1984
Earlier work this paper cites.
A new model-based speech analysis/synthesis system
D Griffin and Jae Lim · 1985
Earlier work this paper cites.
Speech coding based on a multi-layer neural network
Shigeo Morishima, H Harashima, and Y Katayama · 1990
Earlier work this paper cites.
A 2.4 kbit/s melp coder candidate for the new us federal standard
Alan McCree, Kwan Truong, E Bryan George, Thomas P Barnwell, and Vishu Viswanathan · 1996
Earlier work this paper cites.
Statistical theory of quantization
Bernard Widrow, Istvan Kollar, and Ming-Chang Liu · 1996
Earlier work this paper cites.
A review of vector quantization techniques
A Vasuki and PT Vanathi · 2006
Earlier work this paper cites.
Adaptive multi-rate wideband speech codec based on celp algorithm: architectural study, implementation & performance analysis
Dipesh Bhagat, Ninad Bhatt, and Yogeshwar Kosta · 2012
Earlier work this paper cites.
Visqol: The virtual speech quality objective listener
Andrew Hines, Jan Skoglund, Anil Kokaram, and Naomi Harte · 2012
Earlier work this paper cites.
Definition of the opus audio codec
Jean-Marc Valin, Koen Vos, and Timothy Terriberry · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Method for the subjective assessment of intermediate quality level of audio systems
B Series · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Overview of the evs codec architecture
Martin Dietz, Markus Multrus, Vaclav Eksler, Vladimir Malenovsky, Erik Norvell, Harald Pobloth, Lei Miao, Zhe Wang, Lasse Laaksonen, Adriana Vasilache, et al · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Earlier work this paper cites.
End-to-end optimized image compression
Johannes Ballé, Valero Laparra, and Eero P Simoncelli · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Integer networks for data compression with latent-variable models
Johannes Ballé, Nick Johnston, and David Minnen · 2018
Cited alongside, same era.
The challenge of realistic music generation: modelling raw audio at scale
Sander Dieleman, Aaron van den Oord, and Karen Simonyan · 2018
Cited alongside, same era.
Efficient Neural Audio Synthesis
Nal Kalchbrenner et al · 2018
Cited alongside, same era.
Wavenet based low rate speech coding
W Bastiaan Kleijn, Felicia SC Lim, Alejandro Luebs, Jan Skoglund, Florian Stimberg, Quan Wang, and Thomas C Walters · 2018
Cited alongside, same era.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Later among the works it cites.
Single channel voice separation for unknown number of speakers under reverberant and noisy settings
Shlomo E Chazan, Lior Wolf, Eliya Nachmani, and Yossi Adi · 2021
Later among the works it cites.
Global - 2021 forecast highlights - cisco
Cisco · 2021
Later among the works it cites.
Differentiable model compression via pseudo quantization noise
Alexandre Défossez, Yossi Adi, and Gabriel Synnaeve · 2021
Later among the works it cites.
Fsd50k: an open dataset of human-labeled sound events
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2019
Cited alongside, same era.
The mtg-jamendo dataset for automatic music tagging
Dmitry Bogdanov, Minz Won, Philip Tovstogan, Alastair Porter, and Xavier Serra · 2019
Cited alongside, same era.
Music source separation in the waveform domain
Alexandre Défossez, Nicolas Usunier, Léon Bottou, and Francis Bach · 2019
Cited alongside, same era.
Low bit-rate speech coding with vq-vae and a wavenet decoder
Cristina Gârbacea, Aäron van den Oord, Yazhe Li, Felicia SC Lim, Alejandro Luebs, Oriol Vinyals, and Thomas C Walters · 2019
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville · 2019
Cited alongside, same era.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Yi Luo and Nima Mesgarani · 2019
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Later among the works it cites.
Text-free prosody-aware generative spoken language modeling
Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, et al · 2021
Later among the works it cites.
Generative speech coding with predictive variance regularization
W Bastiaan Kleijn, Andrew Storus, Michael Chinen, Tom Denton, Felicia SC Lim, Alejandro Luebs, Jan Skoglund, and Hengchin Yeh · 2021
Later among the works it cites.
Textless speech emotion conversion using decomposed and discrete representations
Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu-Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi · 2021
Later among the works it cites.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al · 2021
Later among the works it cites.
Variable bitrate discrete neural representations via causal self-attention
Shuyang Li, Huanru Henry Mao, and Julian McAuley · 2021
Later among the works it cites.
Real-time speech frequency bandwidth extension
Yunpeng Li, Marco Tagliasacchi, Oleg Rybakov, Victor Ungureanu, and Dominik Roblek · 2021
Later among the works it cites.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Later among the works it cites.
Gan vocoder: Multi-resolution discriminator is all you need
Jaeseong You, Dalhyun Kim, Gyuhyeon Nam, Geumbyeol Hwang, and Gyeongsu Chae · 2021
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Later among the works it cites.
Hifi++: a unified framework for neural vocoding, bandwidth extension and speech enhancement
Pavel Andreev, Aibek Alanov, Oleg Ivanov, and Dmitry Vetrov · 2022
Closest in time.
Icassp 2022 deep noise suppression challenge
Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sergiy Matusevych, Sebastian Braun, Emre Sefik Eskimez, Manthan Thakker, Takuya Yoshioka, Hannes Gamper, and Robert Aichner · 2022
Closest in time.
It’s raw! audio generation with state-space models
Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré · 2022
Closest in time.
Architecture for variable bitrate neural speech codec with configurable computation complexity
Tejas Jayashankar, Thilo Koehler, Kaustubh Kalgaonkar, Zhiping Xiu, Jilong Wu, Ju Lin, Prabhav Agrawal, and Qing He · 2022
Closest in time.
End-to-end neural speech coding for real-time communications
Xue Jiang, Xiulian Peng, Chengyu Zheng, Huaying Xue, Yuan Zhang, and Yan Lu · 2022
Closest in time.
Speech enhancement for low bit rate speech codec
Ju Lin, Kaustubh Kalgaonkar, Qing He, and Xin Lei · 2022
Closest in time.
Generative spoken dialogue language modeling
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al · 2022
Closest in time.
Disentangling speech from surroundings in a neural audio codec
Ahmed Omran, Neil Zeghidour, Zalán Borsos, Félix de Chaumont Quitry, Malcolm Slaney, and Marco Tagliasacchi · 2022
Closest in time.
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee · 2022
Closest in time.