Fetching the paper…
Reading the bibliography…
The recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible.
C. W. J. Granger and M. J. Morris, “Time series modelling and interpretation,” Journal of the Royal Statistical Society: Series A (General) , vol. 139, no. 2, pp. 246–257, 1976
1976
Earlier work this paper cites.
D. O’Shaughnessy, Speech Communications: Human And Machine (IEEE) . Universities press, 1987
1987
Earlier work this paper cites.
Recommendation ITU-T P.800 Methods for subjective determination of transmission quality , ITU-T Std., Aug 1996
1996
Earlier work this paper cites.
J. M. Valin, K. Vos, and T. Terriberry, Definition of the Opus Audio Codec , IETF Std., Sept 2012, RfC: 6717
2012
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Li, K. Swersky, and R. Zemel, “Generative moment matching networks,” in International Conference on Machine Learning , 2015, pp. 1718–1727
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Nowozin, B. Cseke, and R. Tomioka, “f-GAN: Training generative neural samplers using variational divergence minimization,” in Advances in neural information processing systems , 2016, pp. 271–279
2016
Earlier work this paper cites.
J. S. Garofolo, D. Graff, D. Paul, and D. Pallett, “CSR-I (WSJ0) Other,” Harvard Dataverse, Tech. Rep., 2016. [Online]. Available: https://doi.org/10.7910/DVN/ZVU9HF
2016
Cited alongside, same era.
C. Valentini-Botinhao, “Noisy speech database for training speech enhancement algorithms and tts models,” University of Edinburgh. School of Informatics. Centre for Speech Technology Research (CSTR), Tech. Rep., 2016
2016
Cited alongside, same era.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein GAN,” arXiv preprint arXiv:1701.07875 , 2017
2017
Cited alongside, same era.
C.-L. Li, W.-C. Chang, Y. Cheng, Y. Yang, and B. Póczos, “MMD GAN: Towards deeper understanding of moment matching network,” in Advances in Neural Information Processing Systems , 2017, pp. 2203–2213
2017
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in 2019 IEEE Int. Conf. Acoust Speech Signal Processing (ICASSP) . IEEE, 2019, pp. 3617–3621
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Klejsa, P. Hedelin, C. Zhou, R. Fejgin, and L. Villemoes, “High-quality speech coding with sample RNN,” in 2019 IEEE Int. Conf. Acoust Speech Signal Processing (ICASSP) . IEEE, 2019, pp. 7155–7159
2019
Later among the works it cites.
C. Gârbacea, A. van den Oord, Y. Li, F. S. Lim, A. Luebs, O. Vinyals, and T. C. Walters, “Low bit-rate speech coding with VQ-VAE and a WaveNet decoder,” in 2019 IEEE Int. Conf. Acoust Speech Signal Processing (ICASSP) . IEEE, 2019, pp. 735–739
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Fonseca, J. Pons, X. Favory, F. Font, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra, “Freesound datasets: a platform for the creation of open audio datasets,” in Proc. 18th Int. Society Music Information Retrieval Conference (ISMIR 2017) , Suzhou, China, 2017, pp. 486–493
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
A. v. d. Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. Driessche, E. Lockhart, L. Cobo, F. Stimberg et al. , “Parallel WaveNet: Fast high-fidelity speech synthesis,” in International conference on machine learning . PMLR, 2018, pp. 3918–3926
2018
Cited alongside, same era.
2018
Cited alongside, same era.
W. B. Kleijn, F. S. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “WaveNet based low rate speech coding,” in 2018 IEEE Int. Conf. Acoust Speech Signal Processing (ICASSP) . IEEE, 2018, pp. 676–680
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J.-M. Valin and J. Skoglund, “A Real-Time Wideband Neural Vocoder at 1.6kb/s Using LPCNet,” in Proc. Interspeech 2019 , 2019, pp. 3406–3410
2019
Later among the works it cites.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
F. S. Lim, W. B. Kleijn, M. Chinen, and J. Skoglund, “Robust low rate speech coding based on cloned networks and wavenet,” in 2020 IEEE Int. Conf. Acoust Speech Signal Processing (ICASSP) . IEEE, 2020, pp. 6769–6773
2020
Later among the works it cites.
R. Fejgin, J. Klejsa, L. Villemoes, and C. Zhou, “Source coding of audio signals with a generative model,” in 2020 IEEE Int. Conf. Acoust Speech Signal Processing (ICASSP) . IEEE, 2020, pp. 341–345
2020
Later among the works it cites.
S. Sonning, C. Schüldt, H. Erdogan, and S. Wisdom, “Performance study of a convolutional time-domain audio separation network for real-time speech denoising,” in 2020 IEEE Int. Conf. Acoust Speech Signal Processing (ICASSP) . IEEE, 2020, pp. 831–835
2020
Later among the works it cites.