Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations.
“Representation learning with contrastive predictive coding,”
A. v. d. Oord, Y. Li, and O. Vinyals, · 2018
Earlier work this paper cites.
“An unsupervised autoregressive model for speech representation learning,”
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, · 2019
Earlier work this paper cites.
“Self-Supervised Speaker Embeddings,”
T. Stafylakis, J. Rohdin, O. Plchot, P. Mizera, and L. Burget, · 2019
Earlier work this paper cites.
“Similarity of neural network representations revisited,”
S. Kornblith, M. Norouzi, H. Lee, and G. Hinton, · 2019
Earlier work this paper cites.
“Voxceleb: Large-scale speaker verification in the wild,”
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, · 2019
Earlier work this paper cites.
“VoxSRC 2019: The first voxceleb speaker recognition challenge,”
J. S. Chung, A. Nagrani, E. Coto, W. Xie, M. McLaren, D. A. Reynolds, and A. Zisserman, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Earlier work this paper cites.
“ECAPA-TDNN: Emphasized Channel Attention, propagation and aggregation in TDNN based speaker verification,”
B. Desplanques, J. Thienpondt, and K. Demuynck, · 2020
Earlier work this paper cites.
“Understanding self-attention of self-supervised audio transformers,”
S. w. Yang, A. T. Liu, and H. y. Lee, · 2020
Earlier work this paper cites.
“Similarity analysis of contextual word representation models,”
J. Wu, Y. Belinkov, H. Sajjad, N. Durrani, F. Dalvi, and J. Glass, · 2020
Earlier work this paper cites.
“Frame-level phoneme-invariant speaker embedding for text-independent speaker recognition on extremely short utterances,”
N. Tawara, A. Ogawa, T. Iwata, M. Delcroix, and T. Ogawa, · 2020
Earlier work this paper cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, · 2021
Earlier work this paper cites.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
S. w. Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. t. Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. y. Lee, · 2021
Cited alongside, same era.
“Self-supervised text-independent speaker verification using prototypical momentum contrastive learning,”
W. Xia, C. Zhang, C. Weng, M. Yu, and D. Yu, · 2021
Cited alongside, same era.
“Similarity analysis of self-supervised speech representations,”
Y.-A. Chung, Y. Belinkov, and J. Glass, · 2021
Cited alongside, same era.
“Probing acoustic representations for phonetic properties,”
D. Ma, N. Ryant, and M. Liberman, · 2021
Cited alongside, same era.
“Emerging properties in self-supervised vision transformers,”
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, · 2021
Cited alongside, same era.
“C3-DINO: Joint contrastive and non-contrastive self-supervised learning for speaker verification,”
C. Zhang and D. Yu, · 2022
Later among the works it cites.
“Deep versus wide: An analysis of student architectures for task-agnostic knowledge distillation of self-supervised speech models,”
T. Ashihara, T. Moriya, K. Matsuura, and T. Tanaka, · 2022
Later among the works it cites.
“WeSpeaker: A research and production oriented speaker embedding learning toolkit,”
H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y. Deng, and Y. Qian, · 2022
Later among the works it cites.
“SpeechGLUE: How well can self-supervised speech models capture linguistic knowledge?,”
T. Ashihara, T. Moriya, K. Matsuura, T. Tanaka, Y. Ijima, T. Asami, M. Delcroix, and Y. Honma, · 2023
Later among the works it cites.
“A comprehensive study on self-supervised distillation for speaker representation learning,”
Z. Chen, Y. Qian, B. Han, Y. Qian, and M. Zeng, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Do wide and deep networks learn the same things? Uncovering how neural network representations vary with width and depth,”
T. Nguyen, M. Raghu, and S. Kornblith, · 2021
Cited alongside, same era.
“Insights on neural representations for end-to-end speech recognition,”
A. Ollerenshaw, M. A. Jalal, and T. Hain, · 2021
Cited alongside, same era.
“Self-supervised speech representation learning: A review,”
A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe, T. N. Sainath, and S. Watanabe, · 2022
Cited alongside, same era.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, · 2022
Cited alongside, same era.
“Label-efficient self-supervised speaker verification with information maximization and contrastive learning,”
T. Lepage and R. Dehak, · 2022
Cited alongside, same era.
“Non-contrastive self-supervised learning of utterance-level speech representations,”
J. Cho, R. Pappagari, P. Żelasko, L. M. Velazquez, J. Villalba, and N. Dehak, · 2022
Cited alongside, same era.
Later among the works it cites.
“Improving dino-based self-supervised speaker verification with progressive cluster-aware training,”
B. Han, W. Huang, Z. Chen, and Y. Qian, · 2023
Later among the works it cites.
“Pushing the limits of self-supervised speaker verification using regularized distillation framework,”
Y. Chen, S. Zheng, H. Wang, L. Cheng, and Q. Chen, · 2023
Later among the works it cites.
“Comparative layer-wise analysis of self-supervised speech models,”
A. Pasad, B. Shi, and K. Livescu, · 2023
Later among the works it cites.
“Analysing the masked predictive coding training criterion for pre-training a speech representation model,”
H. Yadav, S. Sitaram, and R. R. Shah, · 2023
Later among the works it cites.
“VoxSRC 2022: The fourth voxceleb speaker recognition challenge,”
J. Huh, A. Brown, J.-w. Jung, J. S. Chung, A. Nagrani, D. Garcia-Romero, and A. Zisserman, · 2023
Later among the works it cites.
“Speech self-supervised representation benchmarking: Are we doing it right?,”
S. Zaiem, Y. Kemiche, T. Parcollet, S. Essid, and M. Ravanelli, · 2023
Later among the works it cites.