Fetching the paper…
Reading the bibliography…
In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications.
“Front-End Factor Analysis for Speaker Verification”
Najim Dehak et al · 2011
Earlier work this paper cites.
“Learning both Weights and Connections for Efficient Neural Network”
Song Han, Jeff Pool, John Tran and William Dally · 2015
Earlier work this paper cites.
“The lottery ticket hypothesis: Finding sparse, trainable neural networks”
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
“Speaker embedding extraction with phonetic information”
Yi Liu, Liang He, Jia Liu and Michael Johnson · 2018
Earlier work this paper cites.
“X-vectors: Robust dnn embeddings for speaker recognition”
David Snyder et al · 2018
Earlier work this paper cites.
“An unsupervised autoregressive model for speech representation learning”
Yu-An Chung, Wei-Ning Hsu, Hao Tang and James Glass · 2019
Earlier work this paper cites.
“SNIP: Single-shot Network Pruning based on Connection Sensitivity”
Namhoon Lee, Thalaiyasingam Ajanthan and Philip Torr · 2019
Earlier work this paper cites.
“Introducing phonetic information to speaker embedding for speaker verification”
Yi Liu, Liang He, Jia Liu and Michael. Johnson · 2019
Earlier work this paper cites.
“Effectiveness of Self-Supervised Pre-Training for ASR”
Alexei Baevski and Abdelrahman Mohamed · 2020
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed and Michael Auli · 2020
Earlier work this paper cites.
“Information-Theoretic Understanding of Population Risk Improvement with Model Compression”
Yuheng Bu, Weihao Gao, Shaofeng Zou and Venugopal Veeravalli · 2020
Earlier work this paper cites.
“DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization”
Shaoshi Ling and Yuzong Liu · 2020
Cited alongside, same era.
“Exploring wav2vec 2.0 on speaker verification and language identification”
Zhiyun Fan, Meng Li, Shiyu Zhou and Bo Xu · 2021
Cited alongside, same era.
“Transformer feed-forward layers are key-value memories”
Mor Geva, Roei Schuster, Jonathan Berant and Omer Levy · 2021
Cited alongside, same era.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units”
Wei-Ning Hsu et al · 2021
Cited alongside, same era.
“TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech”
Andy Liu, Shang-Wen Li and Hung-yi Lee · 2021
Cited alongside, same era.
“SUPERB: Speech processing Universal PERformance Benchmark”
Yingzhi Wang, Abdelmoumene Boumadane and Abdelwahab Heba · 2022
Later among the works it cites.
“Phonetic analysis of self-supervised representations of english speech”
Dan Wells, Hao Tang and Korin Richmond · 2022
Later among the works it cites.
“MelHuBERT: A simplified HuBERT on Mel spectrograms”
Tzu-Quan Lin, Hung-yi Lee and Hao Tang · 2023
Later among the works it cites.
“Probing self-supervised speech models for phonetic and phonemic information: a case study in aspiration”
Kinan Martin, Jon Gauthier, Canaan Breiss and Roger Levy · 2023
Later among the works it cites.
“Self-Supervised Speech Representations are More Phonetic than Semantic”
Kwanghee Choi et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shu-wen Yang et al · 2021
Cited alongside, same era.
“data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language”
Alexei Baevski et al · 2022
Cited alongside, same era.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing”
Sanyuan Chen et al · 2022
Cited alongside, same era.
“Large-scale self-supervised speech representation learning for automatic speaker verification”
Zhengyang Chen et al · 2022
Cited alongside, same era.
“Self-supervised pre-training for attention-based encoder-decoder asr model”
Changfeng Gao et al · 2022
Cited alongside, same era.
“Compressing transformer-based self-supervised models for speech processing”
Tzu-Quan Lin et al · 2022
Cited alongside, same era.
“What do end-to-end speech models learn about speaker, language and channel information? A layer-wise and neuron-level analysis”
Shammur Chowdhury, Nadir Durrani and Ahmed Ali · 2024
Later among the works it cites.
“Property Neurons in Self-Supervised Speech Transformers”
Tzu-Quan Lin, Guan-Ting Lin, Hung-yi Lee and Hao Tang · 2024
Later among the works it cites.
“DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning”
Alexander Liu et al · 2024
Later among the works it cites.
“Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models”
Victor Miara, Theo Lepage and Reda Dehak · 2024
Later among the works it cites.
“Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations”
Mukhtar Mohamed, Oli Liu, Hao Tang and Sharon Goldwater · 2024
Later among the works it cites.