Fetching the paper…
Reading the bibliography…
In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing.
“A method of estimating the equal error rate for automatic speaker verification,”
Jyh-Min Cheng and Hsiao-Chuan Wang, · 2004
Earlier work this paper cites.
“Robust feature extraction of speech via noise reduction in autocorrelation domain,”
Gholamreza Farahani, Seyed Mohammad Ahadi, and Mohammad Mehdi Homayounpour, · 2006
Earlier work this paper cites.
“An introduction to application-independent evaluation of speaker recognition systems,”
David A Van Leeuwen and Niko Brümmer, · 2007
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Effective approaches to attention-based neural machine translation,”
Minh-Thang Luong, Hieu Pham, and Christopher D Manning, · 2015
Earlier work this paper cites.
“Musan: A music, speech, and noise corpus,”
David Snyder, Guoguo Chen, and Daniel Povey, · 2015
Earlier work this paper cites.
“A priori snr estimation and noise estimation for speech enhancement,”
Rui Yao, ZeQing Zeng, and Ping Zhu, · 2016
Earlier work this paper cites.
“Long short-term memory-networks for machine reading,”
Jianpeng Cheng, Li Dong, and Mirella Lapata, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Speaker verification with short utterances: a review of challenges, trends and opportunities,”
Arnab Poddar, Md Sahidullah, and Goutam Saha, · 2017
Cited alongside, same era.
“Enhanced feature extraction for speech detection in media audio.,”
Inseon Jang, ChungHyun Ahn, Jeongil Seo, and Younseon Jang, · 2017
Cited alongside, same era.
“Automatic speech emotion recognition using recurrent neural networks with local attention,”
Seyedmahdad Mirsamadi, Emad Barsoum, and Cha Zhang, · 2017
Cited alongside, same era.
“Voxceleb: a large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Cited alongside, same era.
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Squeeze-and-excitation networks,”
Jie Hu, Li Shen, and Gang Sun, · 2018
Later among the works it cites.
“Unified hypersphere embedding for speaker recognition,”
Mahdi Hajibabaei and Dengxin Dai, · 2018
Later among the works it cites.
“Speech enhancement with variational autoencoders and alpha-stable distributions,”
Simon Leglaive, Umut Şimşekli, Antoine Liutkus, Laurent Girin, and Radu Horaud, · 2019
Later among the works it cites.
“Audio-visual speech enhancement using conditional variational auto-encoder,”
Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin, and Radu Horaud, · 2019
Later among the works it cites.
“An adaptive a priori snr estimator for perceptual speech enhancement,”
Lara Nahma, Pei Chee Yong, Hai Huyen Dam, and Sven Nordholm, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Attention-based models for text-dependent speaker verification,”
FA Rezaur rahman Chowdhury, Quan Wang, Ignacio Lopez Moreno, and Li Wan, · 2018
Cited alongside, same era.
“Self-attentive speaker embeddings for text-independent speaker verification.,”
Yingke Zhu, Tom Ko, David Snyder, Brian Mak, and Daniel Povey, · 2018
Cited alongside, same era.
“Attention mechanism in speaker recognition: What does it learn in deep speaker embedding?,”
Qiongqiong Wang, Koji Okabe, Kong Aik Lee, Hitoshi Yamamoto, and Takafumi Koshinaka, · 2018
Cited alongside, same era.
“Attention based fully convolutional network for speech emotion recognition,”
Yuanyuan Zhang, Jun Du, Zirui Wang, Jianshu Zhang, and Yanhui Tu, · 2018
Cited alongside, same era.
“Cbam: Convolutional block attention module,”
Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon, · 2018
Cited alongside, same era.
“Adam: Amethod for stochastic optimization,”
Diederik P Kingma and Jimmy Lei Ba,
Cited in the paper.
Suwon Shon, Hao Tang, and James Glass, · 2019
Later among the works it cites.
“Self multi-head attention for speaker recognition,”
Miquel India, Pooyan Safari, and Javier Hernando, · 2019
Later among the works it cites.
“Deep cnns with self-attention for speaker identification,”
Nguyen Nang An, Nguyen Quang Thanh, and Yanbing Liu, · 2019
Later among the works it cites.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Later among the works it cites.
“Utterance-level aggregation for speaker recognition in the wild,”
Weidi Xie, Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2019
Later among the works it cites.