Fetching the paper…
Reading the bibliography…
Recent work has shown that it is possible to resynthesize high-quality speech based, not on text, but on low bitrate discrete units that have been learned in a self-supervised fashion and can therefore capture expressive aspects of speech that are hard to transcribe (prosody, voice styles, non-verbal vocalization).
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: a resource for the next generations of speech-to-text,” in LREC , 2004
2004
Earlier work this paper cites.
T. Schatz, V. Peddinti, F. Bach, A. Jansen, H. Hermansky, and E. Dupoux, “Evaluating speech features with the Minimal-Pair ABX task: Analysis of the classical MFC/PLP pipeline,” in INTERSPEECH , 2013. [Online]. Available: https://hal.archives-ouvertes.fr/hal-00918599
2013
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in ICASSP , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and K. MacDonald, “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. A. Nguyen, M. de Seyssel, P. Rozé, M. Rivière, E. Kharitonov, A. Baevski, E. Dunbar, and E. Dupoux, “The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,” 2020
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 33, 2020, pp. 12 449–12 460
2020
Earlier work this paper cites.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, “Transformers: State-of-the-art natural language processing,” in Proceedings of Empirical Methods in Natural Language Processing: System Demonstrations , 2020, pp. 38–45. [Online]. Available: https://www.aclweb.org/anthology/2020.emnlp-demos.6
2020
Earlier work this paper cites.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of the Twelfth Language Resources and Evaluation Conference . Marseille, France: European Language Resources Association, May 2020, pp. 4218–4222. [Online]. Available: https://aclanthology.org/2020.lrec-1.520
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP , 2020, pp. 7669–7673, https://github.com/facebookresearch/libri-light
2020
Cited alongside, same era.
A. Clifton, S. Reddy, Y. Yu, A. Pappu, R. Rezapour, H. Bonab, M. Eskevich, G. Jones, J. Karlgren, B. Carterette, and R. Jones, “100,000 podcasts: A spoken English document corpus,” in Proceedings of the 28th International Conference on Computational Linguistics , Dec. 2020, pp. 5903–5917. [Online]. Available: https://aclanthology.org/2020.coling-main.519
2020
Cited alongside, same era.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proceedings of Association for Computational Linguistics , Aug. 2021, pp. 993–1003. [Online]. Available: https://aclanthology.org/2021.acl-long.80
2021
Later among the works it cites.
D. Galvez, G. Diamos, J. Torres, K. Achorn, J. Cerón, A. Gopi, D. Kanter, M. Lam, M. Mazumder, and V. Janapa Reddi, “The people’s speech: A large-scale diverse english speech recognition dataset for commercial usage,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks , J. Vanschoren and S. Yeung, Eds., vol. 1, 2021
2021
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. Audio, Speech, Language Process. (TASLP) , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Trans. Audio, Speech and Lang. Proc. , vol. 30, p. 495–507, nov 2021. [Online]. Available: https://doi.org/10.1109/TASLP.2021.3129994
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Q. Xu, A. Baevski, T. Likhomanenko, P. Tomasello, A. Conneau, R. Collobert, G. Synnaeve, and M. Auli, “Self-training and pre-training are complementary for speech recognition,” in ICASSP . IEEE, 2021, pp. 3030–3034
2021
Cited alongside, same era.
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T. hsien Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. rahman Mohamed, and H. yi Lee, “Superb: Speech processing universal performance benchmark,” in Interspeech , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
E. Kharitonov, J. Copet, K. Lakhotia, T. A. Nguyen, P. Tomasello, A. Lee, A. Elkahky, W.-N. Hsu, A. Mohamed, E. Dupoux, and Y. Adi, “textless-lib: a library for textless spoken language processing,” 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
I. Gat, F. Kreuk, T. A. Nguyen, A. Lee, J. Copet, G. Synnaeve, E. Dupoux, and Y. Adi, “Augmentation invariant discrete representation for generative spoken language modeling,” in Proceedings of the 20th International Conference on Spoken Language Translation (IWSLT 2023) , 2023, pp. 465–477
2023
Closest in time.
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei, “Neural codec language models are zero-shot text to speech synthesizers,” 2023
2023
Closest in time.
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour, “Speak, read and prompt: High-fidelity text-to-speech with minimal supervision,” 2023
2023
Closest in time.