Fetching the paper…
Reading the bibliography…
We introduce a new beam search decoder that is fully differentiable, making it possible to optimize at training time through the inference procedure.
Maximum mutual information estimation of hidden markov model parameters for speech recognition
Lalit Bahl, Peter Brown, Peter De Souza, and Robert Mercer · 1986
Earlier work this paper cites.
Une Approche théorique de l’Apprentissage Connexionniste: Applications à la Reconnaissance de la Parole
Léon Bottou · 1991
Earlier work this paper cites.
The design for the wall street journal-based CSR corpus
Douglas B Paul and Janet M Baker · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Minimum bayes-risk automatic speech recognition
Vaibhava Goel and William J Byrne · 2000
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira · 2001
Earlier work this paper cites.
Incremental parsing with the perceptron algorithm
Michael Collins and Brian Roark · 2004
Earlier work this paper cites.
Graph transformer networks for image recognition
Léon Bottou and Yann LeCun · 2005
Earlier work this paper cites.
Learning as search optimization: Approximate large margin methods for structured prediction
Hal Daumé III and Daniel Marcu · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Hypothesis spaces for minimum bayes risk training in large vocabulary speech recognition
Matthew Gibson and Thomas Hain · 2006
Earlier work this paper cites.
A novel approach to on-line handwriting recognition based on bidirectional long short-term memory networks
Marcus Liwicki, Alex Graves, Horst Bunke, and Jürgen Schmidhuber · 2007
Earlier work this paper cites.
KenLM: faster and smaller language model queries
Kenneth Heafield · 2011
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
Alex Graves and Navdeep Jaitly · 2014
Cited alongside, same era.
Fast and accurate recurrent neural network acoustic models for speech recognition
Haşim Sak, Andrew Senior, Kanishka Rao, and Françoise Beaufays · 2015
Cited alongside, same era.
Sequence level training with recurrent neural networks
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Later among the works it cites.
Letterbased speech recognition with gated convnets
V Liptchinsky, G Synnaeve, and R Collobert · 2017
Later among the works it cites.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles · 2017
Later among the works it cites.
Latent sequence decompositions
William Chan, Yu Zhang, Quoc Le, and Navdeep Jaitly · 2017
Later among the works it cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin · 2017
Later among the works it cites.
Promising accurate prefix boosting for sequence-to-sequence ASR
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush · 2016
Cited alongside, same era.
Globally normalized transition-based neural networks
Daniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, and Michael Collins · 2016
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Cited alongside, same era.
Connectionist temporal modeling for weakly supervised action labeling
De-An Huang, Li Fei-Fei, and Juan Carlos Niebles · 2016
Cited alongside, same era.
Wav2letter: an end-to-end convnet-based speech recognition system
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P. Kingma · 2016
Cited alongside, same era.
Murali Karthick Baskar, Lukáš Burget, Shinji Watanabe, Martin Karafiát, Takaaki Hori, and Jan Honza Černockỳ · 2018
Later among the works it cites.
Fully convolutional speech recognition
Neil Zeghidour, Qiantong Xu, Vitaliy Liptchinsky, Nicolas Usunier, Gabriel Synnaeve, and Ronan Collobert · 2018
Later among the works it cites.
End-to-end speech recognition using lattice-free MMI
Hossein Hadian, Hossein Sameti, Daniel Povey, and Sanjeev Khudanpur · 2018
Later among the works it cites.
Minimum word error rate training for attention-based sequence-to-sequence models
Rohit Prabhavalkar, Tara N Sainath, Yonghui Wu, Patrick Nguyen, Zhifeng Chen, Chung-Cheng Chiu, and Anjuli Kannan · 2018
Later among the works it cites.
Cold fusion: Training seq2seq models together with language models
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates · 2018
Later among the works it cites.
Improving end-to-end speech recognition with policy learning
Yingbo Zhou, Caiming Xiong, and Richard Socher · 2018
Later among the works it cites.
A continuous relaxation of beam search for end-to-end training of neural sequence models
Kartik Goyal, Graham Neubig, Chris Dyer, and Taylor Berg-Kirkpatrick · 2018
Later among the works it cites.
Differentiable dynamic programming for structured prediction and attention
Arthur Mensch and Mathieu Blondel · 2018
Later among the works it cites.