Fetching the paper…
Reading the bibliography…
Speech coding facilitates the transmission of speech over low-bandwidth networks with minimal distortion.
W. B. Kleijn, P. Kroon, and D. Nahumi, “The rcelp speech-coding algorithm,” European Transactions on Telecommunications , vol. 5, no. 5, pp. 573–582, 1994
1994
Earlier work this paper cites.
M. R. Bielefeld and L. M. Supplee, “Developing a test program for the DoD 2400 bps vocoder selection process,” in International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings , vol. 2. IEEE, 1996, pp. 1141–1144
1996
Earlier work this paper cites.
A. McCree, K. Truong, E. B. George, T. P. Barnwell, and V. Viswanathan, “A 2.4 kbit/s MELP coder candidate for the new US federal standard,” in International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings , vol. 1. IEEE, 1996, pp. 200–203
1996
Earlier work this paper cites.
ITU-R, “Recommendation BS.1534-1: Method for the subjective assessment of intermediate quality level of coding systems,” International Telecommunications Union, Geneva, Switzerland , vol. 2, 2001
2001
Earlier work this paper cites.
T. Wang, K. Koishida, V. Cuperman, A. Gersho, and J. Collura, “A 1200/2400 bps coding suite based on MELP,” in Speech Coding, 2002, IEEE Workshop Proceedings. , 2002, pp. 90–92
2002
Earlier work this paper cites.
K. Kasi and S. A. Zahorian, “Yet Another Algorithm for Pitch Tracking,” in International Conference on Acoustics, Speech, and Signal Processing , vol. 1, 2002, pp. I–361–I–364
2002
Earlier work this paper cites.
A. Vasuki and P. Vanathi, “A review of vector quantization techniques,” IEEE Potentials , vol. 25, no. 4, pp. 39–47, 2006
2006
Earlier work this paper cites.
J.-M. Valin, K. Vos, and T. Terriberry, “Definition of the Opus audio codec,” IETF, September , 2012
2012
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
M. Dietz, M. Multrus, V. Eksler, V. Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache et al. , “Overview of the EVS codec architecture,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5698–5702
2015
Earlier work this paper cites.
J.-M. Valin, “Speex: A free codec for free speech,” arXiv preprint arXiv:1602.08668 , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, “End-to-end text-dependent speaker verification,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 5115–5119
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
J. H. Lim and J. C. Ye, “Geometric GAN,” arXiv preprint arXiv:1705.02894 , 2017
2017
Cited alongside, same era.
A. van den Oord, O. Vinyals, and k. kavukcuoglu, “Neural Discrete Representation Learning,” in Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017
2017
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Later among the works it cites.
F. S. Lim, W. B. Kleijn, M. Chinen, and J. Skoglund, “Robust low rate speech coding based on cloned networks and wavenet,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6769–6773
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. B. Kleijn, F. S. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “Wavenet Based Low Rate Speech Coding,” in International conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 676–680
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
J.-M. Valin and J. Skoglund, “LPCNet: Improving Neural Speech Synthesis Through Linear Prediction,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5891–5895
2019
Cited alongside, same era.
J. Klejsa, P. Hedelin, C. Zhou, R. Fejgin, and L. Villemoes, “High-quality speech coding with Sample RNN,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 7155–7159
2019
Cited alongside, same era.
C. Gârbacea, A. van den Oord, Y. Li, F. S. Lim, A. Luebs, O. Vinyals, and T. C. Walters, “Low bit-rate speech coding with VQ-VAE and a WaveNet decoder,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 735–739
2019
Cited alongside, same era.
S. Jafarlou, S. Khorram, V. Kothapally, and J. H. Hansen, “Analyzing large receptive field convolutional networks for distant speech recognition,” in Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 252–259
2019
Cited alongside, same era.
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92),” 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
M. Chinen, F. S. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ViSQOL v3: An open source production ready objective speech and audio metric,” in International conference on quality of multimedia experience (QoMEX) . IEEE, 2020, pp. 1–6
2020
Later among the works it cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
W. A. Jassim, J. Skoglund, M. Chinen, and A. Hines, “WARP-Q: Quality prediction for generative neural speech codecs,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 401–405
2021
Later among the works it cites.