Fetching the paper…
Reading the bibliography…
This paper investigates self-supervised pre-training for audio-visual speaker representation learning where a visual stream showing the speaker's mouth area is used alongside speech as inputs.
Z. N. Karam, W. M. Campbell, and N. Dehak, “Towards reduced false-alarms using cohorts,” ICASSP , pp. 4512–4515, 2011
2011
Earlier work this paper cites.
S. Cumani, P. D. Batzu, D. Colibro, C. Vair, P. Laface, and V. Vasilakakis, “Comparison of speaker recognition approaches for real applications,” in Interspeech , 2011
2011
Earlier work this paper cites.
A. Senior and I. Lopez-Moreno, “Improving DNN speaker independence with i-vector inputs,” in ICASSP , 2014
2014
Earlier work this paper cites.
Y. Lei, N. Scheffer, L. Ferrer, and M. McLaren, “A novel scheme for speaker recognition using a phonetically-aware deep neural network,” in ICASSP , 2014
2014
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” ICLR , 12 2014
2014
Earlier work this paper cites.
J. Hansen and T. Hasan, “Speaker recognition by machines and humans: A tutorial review,” Signal Processing Magazine, IEEE , vol. 32, pp. 74–99, 11 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
T. Stafylakis and G. Tzimiropoulos, “Combining residual networks with LSTMs for lipreading,” in Interspeech , 2017
2017
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Q. Wang, C. Downey, L. Wan, P. A. Mansfield, and I. Lopez-Moreno, “Speaker diarization with lstm,” 2018, pp. 5239–5243
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-Vectors: Robust DNN embeddings for speaker recognition,” in ICASSP , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL , 2018
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep Speaker Recognition,” in Interspeech , 2018
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in ICASSP . IEEE, 2018, pp. 5329–5333
2018
Cited alongside, same era.
A. Nagrani, S. Albanie, and A. Zisserman, “Seeing voices and hearing faces: Cross-modal biometric matching,” in CVPR , 2018
2018
Cited alongside, same era.
Q. Wang et al. , “Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,” in Interspeech , 2019
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Cited alongside, same era.
S. Shon, T.-H. Oh, and J. Glass, “Noise-tolerant audio-visual online person verification using an attention-based neural network fusion,” in ICASSP . IEEE, 2019, pp. 3995–3999
2021
Later among the works it cites.
W. Xia, C. Zhang, C. Weng, M. Yu, and D. Yu, “Self-supervised text-independent speaker verification using prototypical momentum contrastive learning,” in ICASSP , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
W.-N. Hsu et al. , “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” in ICASSP , 2021
2021
Later among the works it cites.
S. Yang et al. , “SUPERB: Speech processing Universal PERformance Benchmark,” in Interspeech , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
L. Zhang, C. Yu, H. Lu, C. Weng, C. Zhang, Y. Wu, X. Xie, Z. Li, and D. Yu, “DurIAN-SC: Duration informed attention network based singing voice conversion system,” in Interspeech , 2020
2020
Cited alongside, same era.
J. S. Chung, J. Huh, S. Mun, M. Lee, H. S. Heo, S. Choe, C. Ham, S. Jung, B.-J. Lee, and I. Han, “In defence of metric learning for speaker recognition,” in Interspeech , 2020
2020
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, propagation and aggregation in TDNN based speaker verification,” in Interspeech , 2020, pp. 3830–3834
2020
Cited alongside, same era.
K. A. Lee, O. Sadjadi, H. Li, and D. Reynolds, “Two decades into speaker recognition evaluation - are we there yet?” Computer Speech and Language , 2020
2020
Cited alongside, same era.
N. Inoue and K. Goto, “Semi-supervised contrastive learning with generalized contrastive loss and its application to speaker recognition,” in APSIPA ASC , 2020
2020
Cited alongside, same era.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS , 2020
2020
Cited alongside, same era.
A. Nagrani, J. S. Chung, S. Albanie, and A. Zisserman, “Disentangled speech embeddings using cross-modal self-supervision,” in ICASSP , 2020, pp. 6829–6833
2020
Cited alongside, same era.
2021
Later among the works it cites.
J. Sarzynska-Wawer et al. , “Detecting formal thought disorder by deep contextualized word representations,” Psychiatry Research , vol. 304, p. 114135, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
L. Sari, K. Singh, J. Zhou, L. Torresani, N. Singhal, and Y. Saraf, “A multi-view approach to audio-visual speaker verification,” 2021
2021
Later among the works it cites.
P. Ma, S. Petridis, and M. Pantic, “End-to-end audio-visual speech recognition with conformers,” in ICASSP , 2021
2021
Later among the works it cites.
J. Thienpondt, B. Desplanques, and K. Demuynck, “The Idlab Voxsrc-20 submission: Large margin fine-tuning and quality-aware score calibration in DNN based speaker verification,” in ICASSP , 2021
2021
Later among the works it cites.
2022
Closest in time.
B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, “Learning audio-visual speech representation by masked multimodal cluster prediction,” 2022
2022
Closest in time.