Fetching the paper…
Reading the bibliography…
Speech tokenization is the task of representing speech signals as a sequence of discrete units.
T. Karrer, E. Lee, and J. O. Borchers, “Phavorit: A phase vocoder for real-time interactive time-stretching,” in ICMC , 2006
2006
Earlier work this paper cites.
L. Yujian and L. Bo, “A normalized levenshtein distance metric,” IEEE transactions on pattern analysis and machine intelligence , vol. 29, no. 6, pp. 1091–1095, 2007
2007
Earlier work this paper cites.
F. Font Corbera, G. Roma Trepat, and X. Serra, “Freesound technical demo,” in MM’13. Proceedings of the 21st ACM international conference on Multimedia; 2013 Oct 21-25; Barcelona, Spain. New York: ACM; 2013. p. 411-2. ACM Association for Computer Machinery, 2013
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “Demand: a collection of multi-channel recordings of acoustic noise in diverse environments,” in Proc. Meetings Acoust , 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP , 2017
2017
Earlier work this paper cites.
R. Scheibler, E. Bezzam, and I. Dokmanić, “Pyroomacoustics: A python package for audio room simulation and array processing algorithms,” in ICASSP , 2018
2018
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” 2019
2019
Earlier work this paper cites.
Y. Adi, N. Zeghidour, R. Collobert, N. Usunier, V. Liptchinsky, and G. Synnaeve, “To reverse the gradient or not: An empirical comparison of adversarial and multi-task learning in speech recognition,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3742–3746
2019
Earlier work this paper cites.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” 2020
2020
Earlier work this paper cites.
T. A. Nguyen, M. de Seyssel, P. Rozé, M. Rivière, E. Kharitonov, A. Baevski, E. Dunbar, and E. Dupoux, “The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,” in NeurIPS – Self-Supervised Learning for Speech and Audio Processing Workshop , 2020
2020
Earlier work this paper cites.
C. K. A. Reddy et al. , “The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,” 2020
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” 2020
2020
Cited alongside, same era.
J. Kahn, M. others Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7669–7673
2020
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Lee et al. , “Textless speech-to-speech translation on real data,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2022, pp. 860–872
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed, and E. Dupoux, “On Generative Spoken Language Modeling from Raw Audio,” TACL , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
S. E. Chazan, L. Wolf, E. Nachmani, and Y. Adi, “Single channel voice separation for unknown number of speakers under reverberant and noisy settings,” in ICASSP , 2021
2021
Cited alongside, same era.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , 2022
2022
Cited alongside, same era.
A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe et al. , “Self-supervised speech representation learning: A review,” IEEE Journal of Selected Topics in Signal Processing , 2022
2022
Cited alongside, same era.
Later among the works it cites.
2022
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Maimon and Y. Adi, “Speaking style conversion in the waveform domain using discrete self-supervised units,” in The 2023 Conference on Empirical Methods in Natural Language Processing , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
I. Gat, F. Kreuk, T. A. Nguyen, A. Lee, J. Copet, G. Synnaeve, E. Dupoux, and Y. Adi, “Augmentation invariant discrete representation for generative spoken language modeling,” in Proceedings of the 20th International Conference on Spoken Language Translation (IWSLT 2023) . Association for Computational Linguistics, 2023, pp. 465–477
2023
Later among the works it cites.