Fetching the paper…
Reading the bibliography…
Sequence-to-sequence models provide a simple and elegant solution for building speech recognition systems by folding separate components of a typical system, namely acoustic (AM), pronunciation (PM) and language (LM) models into a single neural network.
“Development of dialect-specific speech recognizers using adaptation methods,”
Vassilios Diakoloukas, Vassilios Digalakis, Leonardo Neumeyer, and Jaan Kaja, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Towards language independent acoustic modeling,”
William Byrne, Peter Beyerlein, Juan M Huerta, Sanjeev Khudanpur, Bhaskara Marthi, John Morgan, Nino Peterek, Joe Picone, Dimitra Vergyri, and T Wang, · 2000
Earlier work this paper cites.
“Recognizing speech of goats, wolves, sheep and… non-natives,”
Dirk Van Compernolle, · 2001
Earlier work this paper cites.
“Towards universal speech recognition,”
Zhirong Wang, Umut Topkara, Tanja Schultz, and Alex Waibel, · 2002
Earlier work this paper cites.
“Multilingual acoustic modeling using graphemes,”
Stephan Kanthak and Hermann Ney, · 2003
Earlier work this paper cites.
“Learning methods in multilingual speech recognition,”
Hui Lin, Li Deng, Jasha Droppo, Dong Yu, and Alex Acero, · 2008
Earlier work this paper cites.
“A study on multilingual acoustic modeling for large vocabulary ASR,”
Hui Lin, Li Deng, Dong Yu, Yi-fan Gong, Alex Acero, and Chin-Hui Lee, · 2009
Earlier work this paper cites.
“Cross-lingual and multi-stream posterior features for low resource LVCSR systems,”
Samuel Thomas, Sriram Ganapathy, and Hynek Hermansky, · 2010
Earlier work this paper cites.
“Google’s cross-dialect Arabic voice search,”
Fadi Biadsy, Pedro J Moreno, and Martin Jansche, · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Large scale distributed deep networks,”
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al., · 2012
Cited alongside, same era.
“Multilingual acoustic models using distributed deep neural networks,”
Georg Heigold, Vincent Vanhoucke, Alan Senior, Patrick Nguyen, M Ranzato, Matthieu Devin, and Jeffrey Dean, · 2013
Cited alongside, same era.
“Multilingual training of deep neural networks,”
Arnab Ghoshal, Pawel Swietojanski, and Steve Renals, · 2013
Cited alongside, same era.
“Multi-task learning in deep neural networks for improved phoneme recognition,”
Michael L Seltzer and Jasha Droppo, · 2013
Cited alongside, same era.
“Fast speaker adaptation of hybrid NN/HMM model for speech recognition based on discriminative learning of speaker code,”
Ossama Abdel-Hamid and Hui Jiang, · 2013
Cited alongside, same era.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
Haşim Sak, Andrew Senior, Kanishka Rao, and Françoise Beaufays, · 2015
Later among the works it cites.
“Towards acoustic model unification across dialects,”
Mohamed Elfeky, Meysam Bastani, Xavier Velez, Pedro Moreno, and Austin Waters, · 2016
Later among the works it cites.
“On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,”
Liang Lu, Xingxing Zhang, and Steve Renais, · 2016
Later among the works it cites.
“Google’s multilingual neural machine translation system: enabling zero-shot translation,”
Melvin Johnson, Mike Schuster, Quoc V Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al., · 2016
Later among the works it cites.
“Cluster adaptive training for deep neural network based acoustic model,”
Tian Tan, Yanmin Qian, and Kai Yu, · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ngoc Thang Vu, David Imseng, Daniel Povey, Petr Motlicek, Tanja Schultz, and Hervé Bourlard, · 2014
Cited alongside, same era.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Cited alongside, same era.
“Multi-accent deep neural network acoustic model with accent-specific top layer using the KLD-regularized model adaptation,”
Yan Huang, Dong Yu, Chaojun Liu, and Yifan Gong, · 2014
Cited alongside, same era.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2015
Cited alongside, same era.
“Multi-dialectical languages effect on speech recognition: Too much choice can hurt,”
Mohamed Elfeky, Pedro Moreno, and Victor Soto, · 2015
Cited alongside, same era.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Cited alongside, same era.
Later among the works it cites.
“Lower Frame Rate Neural Network Acoustic Models,”
Golan Pundak and Tara N Sainath, · 2016
Later among the works it cites.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,”
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al., · 2016
Later among the works it cites.
“Multi-accent speech recognition with hierarchical grapheme based models,”
Kanishka Rao and Haşim Sak, · 2017
Closest in time.
“Very deep convolutional networks for end-to-end speech recognition,”
Yu Zhang, William Chan, and Navdeep Jaitly, · 2017
Closest in time.
“A Comparison of Sequence-to-Sequence Models for Speech Recognition,”
Rohit Prabhavalkar, Kanishka Rao, Tara N Sainath, Bo Li, Leif Johnson, and Navdeep Jaitly, · 2017
Closest in time.
“Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,”
Chanwoo Kim, Ananya Misra, Kean Chin, Thad Hughes, Arun Narayanan, Tara Sainath, and Michiel Bacchiani, · 2017
Closest in time.