Fetching the paper…
Reading the bibliography…
Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization.
“The htk book,”
Steve Young, Gunnar Evermann, Mark Gales, Thomas Hain, Dan Kershaw, Xunying Liu, Gareth Moore, Julian Odell, Dave Ollason, Dan Povey, et al., · 2002
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“MUSAN: A Music, Speech, and Noise Corpus,” 2015,
David Snyder, Guoguo Chen, and Daniel Povey, · 2015
Earlier work this paper cites.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,”
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al., · 2016
Earlier work this paper cites.
“Sphereface: Deep hypersphere embedding for face recognition,”
Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song, · 2017
Earlier work this paper cites.
“Analysis of score normalization in multilingual speaker recognition.,”
Pavel Matejka, Ondrej Novotnỳ, Oldrich Plchot, Lukas Burget, Mireia Diez Sánchez, and Jan Cernockỳ, · 2017
Earlier work this paper cites.
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Earlier work this paper cites.
“Espnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson-Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al., · 2018
Earlier work this paper cites.
“Unified hypersphere embedding for speaker recognition,”
Mahdi Hajibabaei and Dengxin Dai, · 2018
Earlier work this paper cites.
“Angular softmax for short-duration text-independent speaker verification,”
Zili Huang, Shuai Wang, and Kai Yu, · 2018
Earlier work this paper cites.
“Additive margin softmax for face verification,”
Feng Wang, Jian Cheng, Weiyang Liu, and Haijun Liu, · 2018
Earlier work this paper cites.
“But system description to voxceleb speaker recognition challenge 2019,”
Hossein Zeinali, Shuai Wang, Anna Silnova, Pavel Matějka, and Oldřich Plchot, · 2019
Cited alongside, same era.
“Pytorch: An imperative style, high-performance deep learning library,”
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al., · 2019
Cited alongside, same era.
“Voxsrc 2019: The first voxceleb speaker recognition challenge,”
Joon Son Chung, Arsha Nagrani, Ernesto Coto, Weidi Xie, Mitchell McLaren, Douglas A Reynolds, and Andrew Zisserman, · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Cited alongside, same era.
“High performance i/o for large scale deep learning,”
“Speechbrain: A general-purpose speech toolkit,”
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, et al., · 2021
Later among the works it cites.
“Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,”
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, Fan Yu, Chao Yang, Zhendong Peng, Xiaoyu Chen, Lei Xie, and Xin Lei, · 2021
Later among the works it cites.
“Asv-subtools: Open source toolkit for automatic speaker verification,”
Fuchuan Tong, Miao Zhao, Jianfeng Zhou, Hao Lu, Zheng Li, Lin Li, and Qingyang Hong, · 2021
Later among the works it cites.
“The idlab voxsrc-20 submission: Large margin fine-tuning and quality-aware score calibration in dnn based speaker verification,”
Jenthe Thienpondt, Brecht Desplanques, and Kris Demuynck, · 2021
Later among the works it cites.
“The speakin system for voxceleb speaker recognition challange 2021,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Aizman, Gavin Maltby, and Thomas Breuel, · 2019
Cited alongside, same era.
“Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,”
Xu Xiang, Shuai Wang, Houjun Huang, Yanmin Qian, and Kai Yu, · 2019
Cited alongside, same era.
“Arcface: Additive angular margin loss for deep face recognition,”
Jiankang Deng, Jia Guo, Xue Niannan, and Stefanos Zafeiriou, · 2019
Cited alongside, same era.
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck, · 2020
Cited alongside, same era.
“Voxsrc 2020: The second voxceleb speaker recognition challenge,”
Arsha Nagrani, Joon Son Chung, Jaesung Huh, Andrew Brown, Ernesto Coto, Weidi Xie, Mitchell McLaren, Douglas A Reynolds, and Andrew Zisserman, · 2020
Cited alongside, same era.
“Voxceleb: Large-scale speaker verification in the wild,”
Arsha Nagrani, Joon Son Chung, Weidi Xie, and Andrew Zisserman, · 2020
Cited alongside, same era.
“Spot the conversation: speaker diarisation in the wild,”
Joon Son Chung, Jaesung Huh, Arsha Nagrani, Triantafyllos Afouras, and Andrew Zisserman, · 2020
Cited alongside, same era.
Miao Zhao, Yufeng Ma, Min Liu, and Minqiang Xu, · 2021
Later among the works it cites.
“Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier,”
Silero Team, · 2021
Later among the works it cites.
“Wenet 2.0: More productive end-to-end speech recognition toolkit,”
Binbin Zhang, Di Wu, Zhendong Peng, Xingchen Song, Zhuoyuan Yao, Hang Lv, Lei Xie, Chao Yang, Fuping Pan, and Jianwei Niu, · 2022
Closest in time.
“Voxsrc 2021: The third voxceleb speaker recognition challenge,”
Andrew Brown, Jaesung Huh, Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2022
Closest in time.
“The sjtu x-lance lab system for cnsrc 2022,”
Zhengyang Chen, Bei Liu, Bing Han, Leying Zhang, and Yanmin Qian, · 2022
Closest in time.
“Id r&d system description to voxceleb speaker recognition challenge 2022,”
Rostislav Makarov, Nikita Torgashov, Alexander Alenin, Ivan Yakovlev, and Anton Okhotnikov, · 2022
Closest in time.
“Cn-celeb: multi-genre speaker recognition,”
Lantian Li, Ruiqi Liu, Jiawen Kang, Yue Fan, Hao Cui, Yunqi Cai, Ravichander Vipperla, Thomas Fang Zheng, and Dong Wang, · 2022
Closest in time.