Fetching the paper…
Reading the bibliography…
We present ProsAudit, a benchmark in English to assess structural prosodic knowledge in self-supervised learning (SSL) speech models.
K. E. Silverman, M. E. Beckman, J. F. Pitrelli, M. Ostendorf, C. W. Wightman, P. Price, J. B. Pierrehumbert, and J. Hirschberg, “Tobi: A standard for labeling english prosody.” in ICSLP , vol. 2, 1992, pp. 867–870
1992
Earlier work this paper cites.
M. Ostendorf, P. J. Price, and S. Shattuck-Hufnagel, “The boston university radio news corpus,” Linguistic Data Consortium , pp. 1–19, 1995
1995
Earlier work this paper cites.
A. Cutler, D. Dahan, and W. Van Donselaar, “Prosody in the comprehension of spoken language: A literature review,” Language and speech , vol. 40, no. 2, pp. 141–201, 1997
1997
Earlier work this paper cites.
B. Höhle, R. Bijeljac-Babic, B. Herold, J. Weissenborn, and T. Nazzi, “Language specific prosodic preferences during the first half year of life: Evidence from german and french infants,” Infant Behavior and Development , vol. 32, no. 3, pp. 262–274, 2009
2009
Earlier work this paper cites.
A. D. Endress and M. D. Hauser, “Word segmentation with universal prosodic cues,” Cognitive psychology , vol. 61, no. 2, pp. 177–199, 2010
2010
Earlier work this paper cites.
D. Dahan, “Prosody and language comprehension,” Wiley Interdisciplinary Reviews: Cognitive Science , vol. 6, no. 5, pp. 441–452, 2015
2015
Earlier work this paper cites.
E. Dunbar, X. N. Cao, J. Benjumea, J. Karadayi, M. Bernard, L. Besacier, X. Anguera, and E. Dupoux, “The zero resource speech challenge 2017,” in 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2017, pp. 323–330
2017
Earlier work this paper cites.
2018
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
M. Riviere, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7414–7418
2020
Cited alongside, same era.
T. A. Nguyen, M. de Seyssel, P. Rozé, M. Rivière, E. Kharitonov, A. Baevski, E. Dunbar, and E. Dupoux, “The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,” in NeuRIPS Workshop on Self-Supervised Learning for Speech and Audio Processing , 2020
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” in Proc. Interspeech 2021 , 2021, pp. 1194–1198
2021
Later among the works it cites.
B. Ludusan, M. Morii, Y. Minagawa, and E. Dupoux, “The effect of different information sources on prosodic boundary perception,” JASA Express Letters , vol. 1, no. 11, p. 115203, 2021
2021
Later among the works it cites.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed et al. , “On generative spoken language modeling from raw audio,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1336–1354, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
J. Kahn, M. Riviere, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7669–7673
2020
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
E. Dunbar, M. Bernard, N. Hamilakis, T. A. Nguyen, M. de Seyssel, P. Rozé, M. Rivière, E. Kharitonov, and E. Dupoux, “The Zero Resource Speech Challenge 2021: Spoken Language Modelling,” in Proc. Interspeech 2021 , 2021, pp. 1574–1578
2021
Cited alongside, same era.
E. Kharitonov, A. Lee, A. Polyak, Y. Adi, J. Copet, K. Lakhotia, T. A. Nguyen, M. Riviere, A. Mohamed, E. Dupoux et al. , “Text-free prosody-aware generative spoken language modeling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 8666–8681
2022
Later among the works it cites.
M. Lavechin, M. de Seyssel, H. Titeux, H. Bredin, G. Wisniewski, A. Cristia, and E. Dupoux, “Can statistical learning bootstrap early language acquisition? a modeling investigation,” PsyArXiv preprint PsyArXiv:rx94d , 2022
2022
Later among the works it cites.
G.-T. Lin, C.-L. Feng, W.-P. Huang, Y. Tseng, T.-H. Lin, C.-A. Li, H.-y. Lee, and N. G. Ward, “On the utility of self-supervised models for prosody-related tasks,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 1104–1111
2023
Closest in time.