Fetching the paper…
Reading the bibliography…
Recently, speaker embeddings extracted with deep neural networks became the state-of-the-art method for speaker verification.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Probabilistic linear discriminant analysis for inferences about identity,”
Simon JD Prince and James H Elder, · 2007
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2011
Earlier work this paper cites.
“Language recognition via i-vectors and dimensionality reduction,”
Najim Dehak, Pedro A Torres-Carrasquillo, Douglas Reynolds, and Reda Dehak, · 2011
Earlier work this paper cites.
“Preliminary investigation of Boltzmann machine classifiers for speaker recognition,”
Themos Stafylakis, Patrick Kenny, Mohammed Senoussaoui, and Pierre Dumouchel, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Rectifier nonlinearities improve neural network acoustic models,”
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng, · 2013
Earlier work this paper cites.
“A novel scheme for speaker recognition using a phonetically-aware deep neural network,”
Yun Lei, Nicolas Scheffer, Luciana Ferrer, and Mitchell McLaren, · 2014
Earlier work this paper cites.
“Deep neural networks for extracting Baum-Welch statistics for speaker recognition,”
Patrick Kenny, Vishwa Gupta, Themos Stafylakis, Pierre Ouellet, and Jahangir Alam, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“MUSAN: A music, speech, and noise corpus,”
David Snyder, Guoguo Chen, and Daniel Povey, · 2015
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Cited alongside, same era.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2015
Cited alongside, same era.
“On autoencoders in the i-vector space for speaker recognition,”
Timur Pekhovsky, Sergey Novoselov, Aleksei Sholohov, and Oleg Kudashev, · 2016
Cited alongside, same era.
“Analysis and optimization of bottleneck features for speaker recognition,”
Alicia Lozano-Diez, Anna Silnova, Pavel Matejka, Ondrej Glembek, Oldrich Plchot, Jan Pešán, Lukáš Burget, and Joaquin Gonzalez-Rodriguez, · 2016
Cited alongside, same era.
“Voxceleb: a large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Later among the works it cites.
“Attention-based models for text-dependent speaker verification,”
FA Chowdhury, Quan Wang, Ignacio Lopez Moreno, and Li Wan, · 2017
Later among the works it cites.
“Convolutional neural networks with binaural representations and background subtraction for acoustic scene classification,”
Yoonchang Han and Jeongsoo Park, · 2017
Later among the works it cites.
“Acoustic scene classification using deep convolutional neural network and multiple spectrograms fusion,”
Zheng Weiping, Yi Jiantao, Xing Xiaotao, Liu Xiangtao, and Peng Shaohu, · 2017
Later among the works it cites.
“X-vectors: Robust DNN embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deep neural network-based speaker embeddings for end-to-end speaker verification,”
David Snyder, Pegah Ghahremani, Daniel Povey, Daniel Garcia-Romero, Yishay Carmiel, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“The 2016 speakers in the wild speaker recognition evaluation.,”
Mitchell McLaren, Luciana Ferrer, Diego Castan, and Aaron Lawson, · 2016
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
“HMM-based phrase-independent i-vector extractor for text-dependent speaker verification,”
Hossein Zeinali, Hossein Sameti, and Lukáš Burget, · 2017
Cited alongside, same era.
“Text-dependent speaker verification based on i-vectors, neural networks and hidden Markov models,”
Hossein Zeinali, Hossein Sameti, Lukáš Burget, et al., · 2017
Cited alongside, same era.
“Online signature verification using i-vector representation,”
Hossein Zeinali, Bagher BabaAli, and Hossein Hadian, · 2017
Cited alongside, same era.
“How to train your speaker embeddings extractor,”
Mitchell McLaren, Diego Castan, Mahesh Kumar Nandwana, Luciana Ferrer, and Emre Yılmaz, · 2018
Closest in time.
“On the use of x-vectors for robust speaker recognition,”
Ondřej Novotnỳ, Oldřich Plchot, Pavel Matějka, Ladislav Mošner, and Ondřej Glembek, · 2018
Closest in time.
“Self-attentive speaker embeddings for text-independent speaker verification,”
Yingke Zhu, Tom Ko, David Snyder, Brian Mak, and Daniel Povey, · 2018
Closest in time.
“Attentive statistics pooling for deep speaker embedding,”
Koji Okabe, Takafumi Koshinaka, and Koichi Shinoda, · 2018
Closest in time.
Hossein Zeinali, Lukas Burget, and Jan Cernocky, · 2018
Closest in time.