Fetching the paper…
Reading the bibliography…
In this paper, a hierarchical attention network to generate utterance-level embeddings (H-vectors) for speaker identification is proposed.
“Switchboard cellular part 1 audio,” https://catalog.ldc.upenn.edu/LDC2001S13 , 2001
David Miller David Graff, Kevin Walker, · 2001
Earlier work this paper cites.
“Callhome american english speech,” https://catalog.ldc.upenn.edu/LDC97S42 , 2001
George Zipperlen Alexandra Canavan, David Graff, · 2001
Earlier work this paper cites.
“Visualizing data using t-sne,”
Laurens van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2010
Earlier work this paper cites.
“2008 nist speaker recognition evaluation training set part 1,” https://catalog.ldc.upenn.edu/LDC2011S05 , 2011
NIST Multimodal Information Group, · 2011
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“A unified speaker-dependent speech separation and enhancement system based on deep neural networks,”
Tian Gao, Jun Du, Li Xu, Cong Liu, Li-Rong Dai, and Chin-Hui Lee, · 2015
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Lei Ba, · 2015
Cited alongside, same era.
“Hierarchical attention networks for document classification,”
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy, · 2016
Cited alongside, same era.
“On the use of plda i-vector scoring for clustering short segments,”
Itay Salmun, Irit Opher, and Itshak Lapidot, · 2016
Cited alongside, same era.
“Deep neural network embeddings for text-independent speaker verification.,”
David Snyder, Daniel Garcia-Romero, Daniel Povey, and Sanjeev Khudanpur, · 2017
Cited alongside, same era.
“Speaker diarization using deep neural network embeddings,”
Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, and Alan McCree, · 2017
“Robust and discriminative speaker embedding via intra-class distance variance regularization.,”
Nam Le and Jean-Marc Odobez, · 2018
Later among the works it cites.
“How to train your speaker embeddings extractor,”
ML McLaren, Diego Castan, Mahesh Kumar Nandwana, Luciana Ferrer, and Emre Yilmaz, · 2018
Later among the works it cites.
“Speaker diarization with lstm,”
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno, · 2018
Later among the works it cites.
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“Attention mechanism in speaker recognition: What does it learn in deep speaker embedding?,”
Qiongqiong Wang, Koji Okabe, Kong Aik Lee, Hitoshi Yamamoto, and Takafumi Koshinaka, · 2018
Later among the works it cites.
“Self-attentive speaker embeddings for text-independent speaker verification.,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Spectrum energy based voice activity detection,”
Jing Pang, · 2017
Cited alongside, same era.
“Training utterance-level embedding networks for speaker identification and verification.,”
Heewoong Park, Sukhyun Cho, Kyubyong Park, Namju Kim, and Jonghun Park, · 2018
Cited alongside, same era.
“On deep speaker embeddings for text-independent speaker recognition,”
Sergey Novoselov, Andrey Shulipa, Ivan Kremnev, Alexandr Kozlov, and Vadim Shchemelinin, · 2018
Cited alongside, same era.
Yingke Zhu, Tom Ko, David Snyder, Brian Mak, and Daniel Povey, · 2018
Later among the works it cites.
“Speaker-aware deep denoising autoencoder with embedded speaker identity for speech enhancement,”
Fu-Kai Chuang, Syu-Siang Wang, Jeih-weih Hung, Yu Tsao, and Shih-Hau Fang, · 2019
Closest in time.
“Automatic hierarchical attention neural network for detecting ad,”
Yilin Pan, Bahman Mirheidari, Markus Reuber, Annalena Venneri, Daniel Blackburn, and Heidi Christensen, · 2019
Closest in time.