Fetching the paper…
Reading the bibliography…
We propose using self-supervised discrete representations for the task of speech resynthesis.
B. S. Atal and S. L. Hanauer, “Speech analysis and synthesis by linear prediction of the speech wave,” The journal of the acoustical society of America , vol. 50, no. 2B, pp. 637–655, 1971
1971
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
D. Griffin and J. Lim, “A new model-based speech analysis/synthesis system,” in ICASSP , 1985
1985
Earlier work this paper cites.
A. McCree, K. Truong, E. B. George, T. P. Barnwell, and V. Viswanathan, “A 2.4 kbit/s melp coder candidate for the new us federal standard,” in ICASSP , 1996
1996
Earlier work this paper cites.
K. Kasi and S. A. Zahorian, “Yet another algorithm for pitch tracking,” ICASSP , 2002
2002
Earlier work this paper cites.
T. Nakatani et al. , “A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,” Speech Communication , 2008
2008
Earlier work this paper cites.
W. Chu and A. Alwan, “Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,” ICASSP , 2009
2009
Earlier work this paper cites.
D. Rowe, “Codec 2-open source speech coding at 2400 bits/s and below,” in TAPR and ARRL 30th Digital Communications Conference , 2011, pp. 80–84
2011
Earlier work this paper cites.
J.-M. Valin, K. Vos, and T. Terriberry, “Definition of the opus audio codec,” IETF, September , 2012
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” ICLR , 2014
2014
Earlier work this paper cites.
B. Series, “Method for the subjective assessment of intermediate quality level of audio systems,” International Telecommunication Union Radiocommunication Assembly , 2014
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in ICASSP , 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, “End-to-end text-dependent speaker verification,” in ICASSP , 2016
2016
Earlier work this paper cites.
A. B. L. Larsen et al. , “Autoencoding beyond pixels using a learned similarity metric,” in ICML , 2016
2016
Earlier work this paper cites.
J.-M. Valin, “Speex: A free codec for free speech,” arXiv preprint arXiv:1602.08668 , 2016
2016
Earlier work this paper cites.
J. Ebbers et al. , “Hidden markov model variational autoencoder for acoustic unit discovery,” in INTERSPEECH 2017 , 2017
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals et al. , “Neural discrete representation learning,” in NeurIPS , 2017
2017
Earlier work this paper cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in Advances in Neural Information Processing Systems , 2017
2017
Cited alongside, same era.
A. Vaswani et al. , “Attention is all you need,” in NeurIPS , 2017
2017
Cited alongside, same era.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/
2017
Cited alongside, same era.
C. Veaux et al. , “CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2017
2017
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,” INTERSPEECH , 2017
2017
Cited alongside, same era.
2019
Later among the works it cites.
A. Baevski et al. , “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in ICLR , 2020
2020
Later among the works it cites.
W.-N. Hsu et al. , “Hubert: How much can a bad teacher benefit ASR pre-training?” in NeurIPS Workshop on Self-Supervised Learning for Speech and Audio Processing Workshop , 2020
2020
Later among the works it cites.
J. Kong et al. , “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in NeurIPS , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “Fftnet: A real-time speaker-dependent neural vocoder,” in ICASSP , 2018
2018
Cited alongside, same era.
W. B. Kleijn, F. S. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “Wavenet based low rate speech coding,” in ICASSP , 2018
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” INTERSPEECH , 2018
2018
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised Pre-Training for Speech Recognition,” in INTERSPEECH , 2019
2019
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” ICASSP , 2019
2019
Cited alongside, same era.
J. Devlin et al. , “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
——, “The Zero Resource Speech Challenge 2020: Discovering Discrete Subword and Word Units,” in Proc. Interspeech 2020 , 2020, pp. 4831–4835
2020
Later among the works it cites.
A. Tjandra, S. Sakti, and S. Nakamura, “Transformer VQ-VAE for Unsupervised Unit Discovery and Speech Synthesis: ZeroSpeech 2020 Challenge,” in INTERSPEECH , 2020
2020
Later among the works it cites.
B. van Niekerk, L. Nortje, and H. Kamper, “Vector-Quantized Neural Networks for Acoustic Unit Discovery in the ZeroSpeech 2020 Challenge,” in INTERSPEECH , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Polyak, L. Wolf, and Y. Taigman, “TTS Skins: Speaker Conversion via ASR,” in INTERSPEECH , 2020
2020
Later among the works it cites.
A. Polyak et al. , “Unsupervised Cross-Domain Singing Voice Conversion,” in INTERSPEECH , 2020
2020
Later among the works it cites.
F. S. Lim et al. , “Robust low rate speech coding based on cloned networks and wavenet,” in ICASSP , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Rivière and E. Dupoux, “Towards unsupervised learning of speech features in the wild,” in SLT 2020: IEEE Spoken Language Technology Workshop , 2020
2020
Later among the works it cites.
J. Kahn et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP , 2020
2020
Later among the works it cites.
2021
Closest in time.
——, “High fidelity speech regeneration with application to speech enhancement,” ICASSP , 2021
2021
Closest in time.