Fetching the paper…
Reading the bibliography…
Conventional automatic speech recognition (ASR) typically performs multi-level pattern recognition tasks that map the acoustic speech waveform into a hierarchy of speech units.
“Backpropagation applied to handwritten zip code recognition,”
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel, · 1989
Earlier work this paper cites.
“The design for the Wall Street Journal-based CSR corpus,”
Douglas B. Paul and Janet M. Baker, · 1992
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Probabilistic and bottle-neck features for LVCSR of meetings,”
Frantisek Grézl, Martin Karafiát, Stanislav Kontár, and Jan Cernocky, · 2007
Earlier work this paper cites.
“The application of hidden Markov models in speech recognition,”
Mark Gales and Steve Young, · 2008
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“Convolutive bottleneck network features for LVCSR,”
K. Vesely, M. Karafiát, and F. Grézl, · 2011
Earlier work this paper cites.
“Supervised sequence labelling,”
Alex Graves, · 2012
Earlier work this paper cites.
“End-to-end phoneme sequence recognition using convolutional neural networks,”
Dimitri Palaz, Ronan Collobert, and Mathew Magimai Doss, · 2013
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel Rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
Min Lin, Qiang Chen, and Shuicheng Yan, · 2013
Cited alongside, same era.
“Rectifier nonlinearities improve neural network acoustic models,”
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng, · 2013
Cited alongside, same era.
“Representation learning: A review and new perspectives,”
Yoshua Bengio, Aaron Courville, and Pascal Vincent, · 2013
Cited alongside, same era.
“A scalable approach to using DNN-derived features in GMM-HMM based acoustic modeling for LVCSR,”
Z.J. Yan, Q. Huo, and J. Xu, · 2013
Cited alongside, same era.
“End-to-end continuous speech recognition using attention-based recurrent NN: First results,”
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Convolutional neural networks-based continuous speech recognition using raw speech signal,”
Dimitri Palaz, Mathew Magimai Doss, and Ronan Collobert, · 2015
Later among the works it cites.
“Learning the speech front-end with raw waveform CLDNNs.,”
Tara N Sainath, Ron J Weiss, Andrew W Senior, Kevin W Wilson, and Oriol Vinyals, · 2015
Later among the works it cites.
“Effective approaches to attention-based neural machine translation,”
Minh-Thang Luong, Hieu Pham, and Christopher D Manning, · 2015
Later among the works it cites.
“Cross-lingual transfer learning during supervised training in low resource scenarios.,”
Amit Das and Mark Hasegawa-Johnson, · 2015
Later among the works it cites.
“Acoustic modelling from the signal domain using CNNs,”
Pegah Ghahremani, Vimal Manohar, Daniel Povey, and Sanjeev Khudanpur, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Striving for simplicity: The all convolutional net,”
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller, · 2014
Cited alongside, same era.
“First-pass large vocabulary continuous speech recognition using bi-directional recurrent DNNs,”
Awni Y Hannun, Andrew L Maas, Daniel Jurafsky, and Andrew Y Ng, · 2014
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“How transferable are features in deep neural networks?,”
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson, · 2014
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in English and Mandarin,”
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, JingDong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, et al., · 2016
Later among the works it cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Later among the works it cites.
“Wav2letter: an end-to-end convnet-based speech recognition system,”
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve, · 2016
Later among the works it cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Later among the works it cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Closest in time.