Fetching the paper…
Reading the bibliography…
In this paper, we propose a new loss function called generalized end-to-end (GE2E) loss, which makes the training of speaker verification models more efficient than our previous tuple-based end-to-end (TE2E) loss function.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“A tutorial on text-independent speaker verification,”
Frédéric Bimbot, Jean-François Bonastre, Corinne Fredouille, Guillaume Gravier, Ivan Magrin-Chagnolleau, Sylvain Meignier, Teva Merlin, Javier Ortega-García, Dijana Petrovska-Delacrétaz, and Douglas A Reynolds, · 2004
Earlier work this paper cites.
“An overview of text-independent speaker recognition: From features to supervectors,”
Tomi Kinnunen and Haizhou Li, · 2010
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2011
Earlier work this paper cites.
“Understanding the exploding gradient problem,”
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio, · 2012
Earlier work this paper cites.
“Small-footprint keyword spotting using deep neural networks,”
Guoguo Chen, Carolina Parada, and Georg Heigold, · 2014
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Cited alongside, same era.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling,”
Haşim Sak, Andrew Senior, and Françoise Beaufays, · 2014
Cited alongside, same era.
“Automatic gain control and multi-style training for robust small-footprint keyword spotting with deep neural networks,”
Rohit Prabhavalkar, Raziel Alvarez, Carolina Parada, Preetum Nakkiran, and Tara N Sainath, · 2015
Cited alongside, same era.
“Locally-connected and convolutional neural networks for small footprint speaker recognition,”
Yu-hsin Chen, Ignacio Lopez-Moreno, Tara N Sainath, Mirkó Visontai, Raziel Alvarez, and Carolina Parada, · 2015
Cited alongside, same era.
“Facenet: A unified embedding for face recognition and clustering,”
Florian Schroff, Dmitry Kalenichenko, and James Philbin, · 2015
Cited alongside, same era.
“End-to-end text-dependent speaker verification,”
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer, · 2016
Later among the works it cites.
“Tomato, tomahto. Google Home now supports multiple users,” https://www.blog.google/products/assistant/tomato-tomahto-google-home-now-supports-multiple-users, 2017
Yury Pinsky, · 2017
Closest in time.
“Voice match will allow Google Home to recognize your voice,” https://www.androidheadlines.com/2017/10/voice-match-will-allow-google-home-to-recognize-your-voice.html, 2017
Mihai Matei, · 2017
Closest in time.
“Deep speaker: an end-to-end neural speaker embedding system,”
Chao Li, Xiaokong Ma, Bing Jiang, Xiangang Li, Xuewei Zhang, Xiao Liu, Ying Cao, Ajay Kannan, and Zhenyao Zhu, · 2017
Closest in time.
“End-to-end attention based text-dependent speaker verification,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“The IBM 2016 speaker recognition system,”
Seyed Omid Sadjadi, Sriram Ganapathy, and Jason W. Pelecanos, · 2016
Cited alongside, same era.
Shi-Xiong Zhang, Zhuo Chen, Yong Zhao, Jinyu Li, and Yifan Gong, · 2017
Closest in time.