Fetching the paper…
Reading the bibliography…
We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters.
Statistical Inference for Probabilistic Functions of Finite State Markov Chains
Leonard E. Baum and Ted Petrie · 1966
Earlier work this paper cites.
Continuous Speech Recognition using Multilayer Perceptrons with Hidden Markov Models
Nathaniel Morgan and Herve Bourlard · 1990
Earlier work this paper cites.
Hierarchical Recurrent Neural Networks for Long-Term Dependencies
Salah Hihi and Yoshua Bengio · 1996
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
John Lafferty, Andrew McCallum, and Fernando Pereira · 2001
Earlier work this paper cites.
Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks
Alex Graves, Santiago Fernandez, Faustino Gomez, and Jurgen Schmiduber · 2006
Earlier work this paper cites.
Deep belief networks for phone recognition
Abdel-rahman Mohamed, George E. Dahl, and Geoffrey E. Hinton · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Karafiat Martin, Burget Luka, Eernocky Jan, and Khudanpur Sanjeev · 2010
Earlier work this paper cites.
Large vocabulary continuous speech recognition with context-dependent dbn-hmms
George E. Dahl, Dong Yu, Li Deng, and Alex Acero · 2011
Earlier work this paper cites.
The Kaldi Speech Recognition Toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannenmann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely · 2011
Earlier work this paper cites.
Acoustic modeling using deep belief networks
Abdel-rahman Mohamed, George E. Dahl, and Geoffrey Hinton · 2012
Earlier work this paper cites.
Application of Pretrained Deep Neural Networks to Large Vocabulary Speech Recognition
Navdeep Jaitly, Patrick Nguyen, Andrew W. Senior, and Vincent Vanhoucke · 2012
Earlier work this paper cites.
Sequence Transduction with Recurrent Neural Networks
Alex Graves · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Cited alongside, same era.
Large Scale Distributed Deep Networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Cited alongside, same era.
Deep Convolutional Neural Networks for LVCSR
Tara Sainath, Abdel-rahman Mohamed, Brian Kingsbury, and Bhuvana Ramabhadran · 2013
Cited alongside, same era.
Hybrid Speech Recognition with Bidirectional LSTM
Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed · 2013
Cited alongside, same era.
Distributed Representations of Words and Phrases and their Compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Cited alongside, same era.
Towards End-to-End Speech Recognition with Recurrent Neural Networks
Grapheme-to-phoneme conversion using long short-term memory recurrent neural networks
Kanishka Rao, Fuchun Peng, Hasim Sak, and Francoise Beaufays · 2015
Closest in time.
Sequence-to-Sequence Neural Net Models for Grapheme-to-Phoneme Conversion
Kaisheng Yao and Geoffrey Zweig · 2015
Closest in time.
Attention-Based Models for Speech Recognition
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Closest in time.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Closest in time.
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Graves and Navdeep Jaitly · 2014
Cited alongside, same era.
Deep Speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and Andrew Ng · 2014
Cited alongside, same era.
End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, and Quoc Le · 2014
Cited alongside, same era.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwen, and Yoshua Bengio · 2014
Cited alongside, same era.
Oriol Vinyals, Lukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey E. Hinton · 2014
Cited alongside, same era.
A Clockwork RNN
Jan Koutnik, Klaus Greff, Faustino Gomez, and Jurgen Schmidhuber · 2014
Cited alongside, same era.
Convolutional, Long Short-Term Memory, Fully Connected Deep Neural Networks
Tara N. Sainath, Oriol Vinyals, Andrew Senior, and Hasim Sak · 2015
Closest in time.
Addressing the Rare Word Problem in Neural Machine Translation
Minh-Thang Luong, Ilya Sutskever, Quoc V. Le, Oriol Vinyals, and Wojciech Zaremba · 2015
Closest in time.
On Using Very Large Target Vocabulary for Neural Machine Translation
Sebastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio · 2015
Closest in time.
Show and Tell: A Neural Image Caption Generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Closest in time.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Closest in time.
A Neural Conversational Model
Oriol Vinyals and Quoc V. Le · 2015
Closest in time.
Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition
Hasim Sak, Andrew Senior, Kanishka Rao, and Francoise Beaufays · 2015
Closest in time.