Fetching the paper…
Reading the bibliography…
We provide a speech coding scheme employing a generative model based on SampleRNN that, while operating at significantly lower bitrates, matches or surpasses the perceptual quality of state-of-the-art classic wide-band codecs.
“Voice conversion with conditional samplernn,”
Cong Zhou, Michael Horgan, Vivek Kumar, Cristina Vasco, and Dan Darcy, · 1977
Earlier work this paper cites.
“Efficient vector quantization of LPC parameters at 24 bits/frame,”
Kuldip K Paliwal and Bishnu S Atal, · 1990
Earlier work this paper cites.
“An efficient gradient-based algorithm for on-line training of recurrent network trajectories,”
Ronald J Williams and Jing Peng, · 1990
Earlier work this paper cites.
“CSR-I (WSJ0) complete LDC93S6A,” Philadelphia: Linguistic Data Consortium, 1993
John S. Garofolo and et al., · 1993
Earlier work this paper cites.
“Improving predictive vector quantizers in speech coding applications,”
Jan Lindén, Jan Skoglund, and Thomas Eriksson, · 1996
Earlier work this paper cites.
“A sinusoidal LPC vocoder,”
Per Hedelin, · 2000
Earlier work this paper cites.
“Spanish statistical parametric speech synthesis using a neural vocoder,”
Antonio Bonafonte, Santiago Pascual, and Georgina Dorca, · 2001
Earlier work this paper cites.
“The adaptive multirate wideband speech codec (AMR-WB),”
B. Bessette, R. Salami, R. Lefebvre, M. Jelinek, J. Rotola-Pukkila, J. Vainio, H. Mikkola, and K. Jarvinen, · 2002
Earlier work this paper cites.
“TSP Speech Database,” Telecommunications & Signal Processing Lab, Dept. of Elec. & Computer Engineering, McGill University, Sept. 2002
Peter Kabal, · 2002
Earlier work this paper cites.
“Entropy constrained quantization of LSP parameters,”
Turaj Zakizadeh Shabestary, Per Hedelin, and Fredrik Nordén, · 2003
Earlier work this paper cites.
“Vector quantization by companding a union of z-lattices,”
Turaj Zakizadeh Shabestary and Per Hedelin, · 2005
Cited alongside, same era.
“SILK speech codec,”
Soeren Jensen, Koen Vos, and Karsten Soerensen, · 2010
Cited alongside, same era.
“Voice quality evaluation of recent open source codecs,”
Anssi Rämö and Henri Toukomaa, · 2010
Cited alongside, same era.
“Perceptual objective listening quality assessment,” Recommendation ITU-T P.863, Jan. 2011
2011
Cited alongside, same era.
Vector quantization and signal compression
Allen Gersho and Robert M Gray, · 2012
Cited alongside, same era.
Training recurrent neural networks
Ilya Sutskever, · 2013
Cited alongside, same era.
“SampleRNN: An unconditional end-to-end neural audio generation model,”
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio, · 2016
Later among the works it cites.
“Nuts and bolts of deep learning,”
Andrew Ng, · 2016
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville, · 2016
Later among the works it cites.
“Deep voice: Real-time neural text-to-speech,”
Sercan O Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al., · 2017
Later among the works it cites.
“Char2wav: End-to-end speech synthesis,”
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Cited alongside, same era.
“ADAM: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Method for the subjective assessment of intermediate quality levels of coding systems,” Recommendation ITU-R BS.1534-3, Oct. 2015
2015
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu, · 2016
Cited alongside, same era.
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P Kingma, · 2017
Later among the works it cites.
“Noisy speech database for training speech enhancement algorithms and TTS models,” University of Edinburgh. School of Informatics. Centre for Speech Technology Research (CSTR), 2017
Cassia Valentini-Botinhao et al., · 2017
Later among the works it cites.
“Parallel WaveNet: Fast high-fidelity speech synthesis,”
Aaron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C Cobo, Florian Stimberg, et al., · 2017
Later among the works it cites.
“Wavenet based low rate speech coding,”
W. B. Kleijn, F. S. C. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, · 2018
Closest in time.