Fetching the paper…
Reading the bibliography…
Attention-based models have recently shown great performance on a range of tasks, such as speech recognition, machine translation, and image captioning due to their ability to summarize relevant information that expands through the entire length of an input sequence.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2011
Earlier work this paper cites.
“Analysis of i-vector length normalization in speaker recognition systems.,”
Daniel Garcia-Romero and Carol Y Espy-Wilson, · 2011
Earlier work this paper cites.
“Small-footprint keyword spotting using deep neural networks,”
Guoguo Chen, Carolina Parada, and Georg Heigold, · 2014
Earlier work this paper cites.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling,”
Haşim Sak, Andrew Senior, and Françoise Beaufays, · 2014
Earlier work this paper cites.
“Deep learning,”
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, · 2015
Earlier work this paper cites.
“Automatic gain control and multi-style training for robust small-footprint keyword spotting with deep neural networks,”
Rohit Prabhavalkar, Raziel Alvarez, Carolina Parada, Preetum Nakkiran, and Tara N Sainath, · 2015
Cited alongside, same era.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Cited alongside, same era.
“Effective approaches to attention-based neural machine translation,”
Minh-Thang Luong, Hieu Pham, and Christopher D Manning, · 2015
Cited alongside, same era.
“Facenet: A unified embedding for face recognition and clustering,”
Florian Schroff, Dmitry Kalenichenko, and James Philbin, · 2015
Cited alongside, same era.
“End-to-end text-dependent speaker verification,”
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer, · 2016
Cited alongside, same era.
“Voice match will allow google home to recognize your voice,” https://www.androidheadlines.com/2017/10/voice-match-will-allow-google-home-to-recognize-your-voice.html, 2017
Mihai Matei, · 2017
Closest in time.
“End-to-end dnn based speaker recognition inspired by i-vector and plda,”
Johan Rohdin, Anna Silnova, Mireia Diez, Oldrich Plchot, Pavel Matejka, and Lukas Burget, · 2017
Closest in time.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2017
Closest in time.
“Speaker diarization with lstm,”
Quan Wang, Carlton Downey, Li Wan, Philip Mansfield, and Ignacio Lopez Moreno, · 2017
Closest in time.
“Show, attend and tell: Neural image caption generation with visual attention,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Tomato, tomahto. google home now supports multiple users,” https://www.blog.google/products/assistant/tomato-tomahto-google-home-now-supports-multiple-users, 2017
Yury Pinsky, · 2017
Cited alongside, same era.
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio, · 2057
Closest in time.