Fetching the paper…
Reading the bibliography…
Self-supervised learning methods such as wav2vec 2.0 have shown promising results in learning speech representations from unlabelled and untranscribed speech data that are useful for speech recognition.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Joint face detection and facial expression recognition with mtcnn,”
Jia Xiang and Gengming Zhu, · 2017
Earlier work this paper cites.
“Focal loss for dense object detection,”
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár, · 2017
Earlier work this paper cites.
“Voxceleb: a large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Earlier work this paper cites.
“Sphereface: Deep hypersphere embedding for face recognition,”
Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song, · 2017
Earlier work this paper cites.
Weicheng Cai, Jinkun Chen, and Ming Li, · 2018
Earlier work this paper cites.
“Speech commands: A dataset for limited-vocabulary speech recognition,”
Pete Warden, · 2018
Earlier work this paper cites.
“Attentive statistics pooling for deep speaker embedding,”
Koji Okabe, Takafumi Koshinaka, and Koichi Shinoda, · 2018
Cited alongside, same era.
“Utterance-level aggregation for speaker recognition in the wild,”
Weidi Xie, Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2019
Cited alongside, same era.
“Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,”
Xu Xiang, Shuai Wang, Houjun Huang, Yanmin Qian, and Kai Yu, · 2019
Cited alongside, same era.
“Improving multi-scale aggregation using feature pyramid module for robust speaker verification of variable-duration utterances,”
Youngmoon Jung, Seong Min Kye, Yeunju Choi, Myunghun Jung, and Hoi-Rin Kim, · 2020
Cited alongside, same era.
“Speech emotion recognition using convolutional neural network and long-short term memory,”
Ranjana Dangol, Abeer Alsadoon, PWC Prasad, Indra Seher, and Omar Hisham Alsadoon, · 2020
Cited alongside, same era.
“Voxceleb: Large-scale speaker verification in the wild,”
Arsha Nagrani, Joon Son Chung, Weidi Xie, and Andrew Zisserman, · 2020
Later among the works it cites.
“Exploring wav2vec 2.0 on speaker verification and language identification,”
Zhiyun Fan, Meng Li, Shiyu Zhou, and Bo Xu, · 2020
Later among the works it cites.
“Broadcasted residual learning for efficient keyword spotting,”
Byeonggeun Kim, Simyung Chang, Jinkyu Lee, and Dooyong Sung, · 2021
Closest in time.
“Multi-channel spectrograms for speech processing applications using deep learning methods,”
Tomas Arias-Vergara, Philipp Klumpp, Juan Camilo Vasquez-Correa, Elmar Nöth, Juan Rafael Orozco-Arroyave, and Maria Schuster, · 2021
Closest in time.
“Joint face image restoration and frontalization for recognition,”
Xiaoguang Tu, Jian Zhao, Qiankun Liu, Wenjie Ai, Guodong Guo, Zhifeng Li, Wei Liu, and Jiashi Feng, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Retinaface: Single-shot multi-level face localisation in the wild,”
Jiankang Deng, Jia Guo, Evangelos Ververas, Irene Kotsia, and Stefanos Zafeiriou, · 2020
Cited alongside, same era.
“Wav2kws: Transfer learning from speech representations for keyword spotting,”
Deokjin Seo, Heung-Seon Oh, and Yuchul Jung, · 2021
Closest in time.
“Emotion recognition from speech using wav2vec 2.0 embeddings,”
Leonardo Pepino, Pablo Riera, and Luciana Ferrer, · 2021
Closest in time.