Fetching the paper…
Reading the bibliography…
End-to-end automatic speech recognition (ASR) commonly transcribes audio signals into sequences of characters while its performance is evaluated by measuring the word-error rate (WER).
Large vocabulary continuous speech recognition using HTK
P. C. Woodland, J. J. Odell, V. Valtchev, and S. J. Young · 1994
Earlier work this paper cites.
Multitask Learning
R. Caruana · 1997
Earlier work this paper cites.
A Model of Inductive Bias Learning
J. Baxter · 2000
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Sequence Labelling in Structured Domains with Hierarchical Recurrent Neural Networks
S. Fernández, A. Graves, and J. Schmidhuber · 2007
Earlier work this paper cites.
Sequence Transduction with Recurrent Neural Networks
A. Graves · 2012
Earlier work this paper cites.
Deep Speech: Scaling up end-to-end speech recognition
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Y. Ng · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, E. Elsen, J. Engel, L. Fan, C. Fougner, A. Y. Hannun, B. Jun, T. Han, P. LeGresley, X. Li, L. Lin, S. Narang, A. Y. Ng, S. Ozair, R. Prenger, S. Qian, J. Raiman, S. Satheesh, D. Seetapun, S. Sengupta, C. Wang, Y. Wang, Z. Wang, B. Xiao, Y. Xie, D. Yogatama, J. Zhan, and Z. Zhu · 2016
Earlier work this paper cites.
End-to-end attention-based large vocabulary speech recognition
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals · 2016
Earlier work this paper cites.
Wav2Letter: An End-to-End ConvNet-based Speech Recognition System
R. Collobert, C. Puhrsch, and G. Synnaeve · 2016
Cited alongside, same era.
Deep multi-task learning with low level tasks supervised at lower layers
A. Søgaard and Y. Goldberg · 2016
Cited alongside, same era.
Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks
Y. Zhang, M. Pezeshki, P. Brakel, S. Zhang, C. Laurent, Y. Bengio, and A. C. Courville · 2016
Cited alongside, same era.
A Closer Look at Memorization in Deep Networks
D. Arpit, S. Jastrzębski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, and S. Lacoste-Julien · 2017
Cited alongside, same era.
Direct Acoustics-to-Word Models for English Conversational Speech Recognition
K. Audhkhasi, B. Ramabhadran, G. Saon, M. Picheny, and D. Nahamoo · 2017
Cited alongside, same era.
Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition
H. Soltau, H. Liao, and H. Sak · 2017
Later among the works it cites.
Multitask Learning with Low-Level Auxiliary Tasks for Encoder-Decoder Based Speech Recognition
S. Toshniwal, H. Tang, L. Lu, and K. Livescu · 2017
Later among the works it cites.
Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition
K. Audhkhasi, B. Kingsbury, B. Ramabhadran, G. Saon, and M. Picheny · 2018
Closest in time.
Multi-Task Learning of Pairwise Sequence Classification Tasks over Disparate Label Spaces
I. Augenstein, S. Ruder, and A. Søgaard · 2018
Closest in time.
GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks
Z. Chen, V. Badrinarayanan, C.-Y. Lee, and A. Rabinovich · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Identifying beneficial task relations for multi-task learning in deep neural networks
J. Bingel and A. Søgaard · 2017
Cited alongside, same era.
Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
A. Kendall, Y. Gal, and R. Cipolla · 2017
Cited alongside, same era.
Joint CTC-attention based end-to-end speech recognition using multi-task learning
S. Kim, T. Hori, and S. Watanabe · 2017
Cited alongside, same era.
Acoustic-to-word model without OOV
J. Li, G. Ye, R. Zhao, J. Droppo, and Y. Gong · 2017
Cited alongside, same era.
Gram-CTC: Automatic Unit Selection and Target Decomposition for Sequence Labelling
H. Liu, Z. Zhu, X. Li, and S. Satheesh · 2017
Cited alongside, same era.
Optimizing Expected Word Error Rate via Sampling for Speech Recognition
M. Shannon · 2017
Cited alongside, same era.
Advancing Connectionist Temporal Classification with Attention Modeling
A. Das, J. Li, R. Zhao, and Y. Gong · 2018
Closest in time.
Hierarchical Multitask Learning for CTC-based Speech Recognition
K. Krishna, S. Toshniwal, and K. Livescu · 2018
Closest in time.
Advancing Acoustic-to-Word CTC Model
J. Li, G. Ye, A. Das, R. Zhao, and Y. Gong · 2018
Closest in time.
Hierarchical Multi Task Learning With CTC
R. Sanabria and F. Metze · 2018
Closest in time.
Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based Model
S. Ueno, H. Inaguma, M. Mimura, and T. Kawahara · 2018
Closest in time.
Improving End-to-End Speech Recognition with Policy Learning
Y. Zhou, C. Xiong, and R. Socher · 2018
Closest in time.