Fetching the paper…
Reading the bibliography…
Human speech data comprises a rich set of domain factors such as accent, syntactic and semantic variety, or acoustic environment.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in Acoustics, Speech, and Signal Processing, IEEE International Conference on , 1992
1992
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: A resource for the next generations of speech-to-text.” in LREC , 2004
2004
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in ICML , 2006
2006
Earlier work this paper cites.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2017
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An unsupervised autoregressive model for speech representation learning,” in Interspeech , 2019
2019
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” in Interspeech , 2019
2019
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in Proc. of NAACL System Demonstrations , 2019
2019
Earlier work this paper cites.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS , 2020
2020
Cited alongside, same era.
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Unsupervised cross-lingual representation learning for speech recognition,” in Interspeech , 2020
2020
Cited alongside, same era.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP , 2020
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in NeurIPS , 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
C. Wang, W.-N. Hsu, Y. Adi, A. Polyak, A. Lee, P.-J. Chen, J. Gu, and J. Pino, “fairseq sˆ2: A scalable and integrable speech synthesis toolkit,” in EMNLP: System Demonstrations , 2021
2021
Later among the works it cites.
“Weights used for synthesize speech.” https://github.com/pytorch/fairseq/tree/main/examples/speech_synthesis , 2021
2021
Later among the works it cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in ICLR , 2021
2021
Later among the works it cites.
I. Papadimitriou and D. Jurafsky, “Learning Music Helps You Read: Using transfer to study linguistic structure in language models,” in EMNLP . Association for Computational Linguistics, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
S.-w. Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin et al. , “Superb: Speech processing universal performance benchmark,” in Interspeech , 2021
2021
Cited alongside, same era.
A. Polyak, Y. Adi, J. Copet, E. Kharitonov, K. Lakhotia, W.-N. Hsu, A. Mohamed, and E. Dupoux, “Speech resynthesis from discrete disentangled self-supervised representations,” in Interspeech , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
“Montreal Forced Aligner,” https://github.com/MontrealCorpusTools/Montreal-Forced-Aligner
Cited in the paper.
2021
Later among the works it cites.
K. Krishna, J. Bigham, and Z. C. Lipton, “Does pretraining for summarization require knowledge transfer?” in EMNLP , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
D. Berrebbi, R. Collobert, N. Jaitly, and T. Likhomanenko, “More speaking or more speakers?” ICASSP , 2023
2023
Closest in time.