Fetching the paper…
Reading the bibliography…
We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs.
J. MacQueen, “Some methods for classification and analysis of multivariate observations,” Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability , pp. 281–297, 1967
1967
Earlier work this paper cites.
Y. Linde, A. Buzo, and R. Gray, “An algorithm for vector quantizer design,” IEEE Transactions on Communications , vol. 28, pp. 84–95, 1980
1980
Earlier work this paper cites.
S. Lloyd, “Least squares quantization in PCM,” IEEE transactions on information theory , vol. 28, pp. 129–137, 1982
1982
Earlier work this paper cites.
B.-H. Juang and A. Gray, “Multiple stage vector quantization for speech coding,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 1982, pp. 597–600
1982
Earlier work this paper cites.
R. Gray, “Vector quantization,” IEEE ASSP Magazine , vol. 1, pp. 4–29, 1984
1984
Earlier work this paper cites.
J. Makhoul, S. Roucos, and H. Gish, “Vector quantization in speech coding,” Proceedings of the IEEE , vol. 73, pp. 1551–1588, 1985
1985
Earlier work this paper cites.
M. Schroeder and B. Atal, “Code-excited linear prediction (CELP): High-quality speech at very low bit rates,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 1985, pp. 937–940
1985
Earlier work this paper cites.
S. Morishima, H. Harashima, and Y. Katayama, “Speech coding based on a multi-layer neural network,” in IEEE International Conference on Communications, Including Supercomm Technical Sessions , 1990, pp. 429–433
1990
Earlier work this paper cites.
ITU-R, Recommendation BS.1534-1: Method for the subjective assessment of intermediate quality level of coding systems , International Telecommunications Union, 2001
2001
Earlier work this paper cites.
ITU, “Perceptual evaluation of speech quality (PESQ): an objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” Int. Telecomm. Union, Geneva, Switzerland, ITU-T Rec. P.862, 2001
2001
Earlier work this paper cites.
B. Bessette, R. Salami, R. Lefebvre, M. Jelinek, J. Rotola-Pukkila, J. Vainio, H. Mikkola, and K. Jarvinen, “The adaptive multirate wideband speech codec (AMR-WB),” IEEE Transactions on Speech and Audio Processing , vol. 10, pp. 620–636, 2002
2002
Earlier work this paper cites.
A. Vasuki and P. Vanathi, “A review of vector quantization techniques,” IEEE Potentials , vol. 25, pp. 39–47, 2006
2006
Earlier work this paper cites.
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of algorithms using games: The case of music tagging.” in ISMIR , 2009, pp. 387–392
2009
Earlier work this paper cites.
J.-M. Valin, K. Vos, and T. B. Terriberry, “Definition of the Opus Audio Codec,” IETF RFC 6716, 2012, https://tools.ietf.org/html/rfc6716
2012
Earlier work this paper cites.
A. Hines, J. Skoglund, A. Kokaram, and N. Harte, “ViSQOL: The virtual speech quality objective listener,” in International Workshop on Acoustic Signal Enhancement (IWAENC) , 2012, pp. 1–4
2012
Earlier work this paper cites.
T. Ishii, H. Komiyama, T. Shinozaki, Y. Horiuchi, and S. Kuroiwa, “Reverberant speech recognition based on denoising autoencoder.” in Interspeech , 2013, pp. 3512–3516
2013
Earlier work this paper cites.
X. Feng, Y. Zhang, and J. Glass, “Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 1759–1763
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research , vol. 15, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
M. Dietz, M. Multrus, V. Eksler, V. Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache, Y. Kamamoto, K. Kikuiri, S. Ragot, J. Faure, H. Ehara, V. Rajendran, V. Atti, H. Sung, E. Oh, H. Yuan, and C. Zhu, “Overview of the EVS codec architecture,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5698–5702
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Holmberg, S. Håkansson, and G. Eriksson, “Web real-time communication use cases and requirements,” IETF RFC 7478, Mar. 2015, https://tools.ietf.org/html/rfc7478
2015
Earlier work this paper cites.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
D. S. Williamson and D. Wang, “Time-frequency masking in the complex domain for speech dereverberation and denoising,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, pp. 1492–1501, 2017
2017
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brebisson, Y. Bengio, and A. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” in Advances in Neural Information Processing Systems , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
E. Fonseca, J. Pons Puig, X. Favory, F. Font Corbera, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra, “Freesound datasets: a platform for the creation of open audio datasets,” in Proceedings of the 18th ISMIR Conference , 2017, pp. 486–493
2017
Cited alongside, same era.
W. B. Kleijn, F. S. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “Wavenet based low rate speech coding,” in IEEE international conference on acoustics, speech and signal processing (ICASSP) , 2018, pp. 676–680
2018
Cited alongside, same era.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, “Parallel WaveNet: Fast high-fidelity speech synthesis,” in Proceedings of the 35th International Conference on Machine Learning , 2018, pp. 3918–3926
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “FFTNet: a real-time speaker-dependent neural vocoder,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 2251–2255
2018
Cited alongside, same era.
2018
Cited alongside, same era.
D. Rethage, J. Pons, and X. Serra, “A WaveNet for speech denoising,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5069–5073
2018
Cited alongside, same era.
W3C, “WebRTC 1.0: Real-time communication between browsers,” 2019, https://www.w3.org/TR/webrtc/
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Biswas and D. Jia, “Audio codec enhancement with generative adversarial networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 356–360
2020
Later among the works it cites.
F. Stimberg, A. Narest, A. Bazzica, L. Kolmodin, P. Barrera González, O. Sharonova, H. Lundin, and T. C. Walters, “WaveNetEQ — Packet loss concealment with WaveRNN,” in 54th Asilomar Conference on Signals, Systems, and Computers , 2020, pp. 672–676
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Tagliasacchi, Y. Li, K. Misiunas, and D. Roblek, “SEANet: A multi-modal speech enhancement network,” in Interspeech , 2020, pp. 1126–1130
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ViSQOL v3: an open source production ready objective speech and audio metric,” in Twelfth International Conference on Quality of Multimedia Experience (QoMEX) , 2020, pp. 1–6
2020
Later among the works it cites.
2020
Later among the works it cites.
Y. Li, M. Tagliasacchi, O. Rybakov, V. Ungureanu, and D. Roblek, “Real-time speech frequency bandwidth extension,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 691–695
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
J. Casebeer, V. Vale, U. Isik, J.-M. Valin, R. Giri, and A. Krishnaswamy, “Enhancing into the codec: Noise robust speech coding with vector-quantized autoencoders,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 711–715
2021
Closest in time.