Fetching the paper…
Reading the bibliography…
Connectionist temporal classification (CTC) is widely used for maximum likelihood learning in end-to-end speech recognition models.
“Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
R J Williams, · 1992
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A Graves, S Fernández, F Gomez, and J Schmidhuber, · 2006
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
G Hinton, L Deng, D Yu, G E Dahl, A Mohamed, N Jaitly, A Senior, V Vanhoucke, P Nguyen, T Sainath, et al., · 2012
Earlier work this paper cites.
“Rotation, scaling and deformation invariant scattering for texture discrimination,”
L Sifre and S Mallat, · 2013
Earlier work this paper cites.
“On the difficulty of training recurrent neural networks,”
R Pascanu, T Mikolov, and Y Bengio, · 2013
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A Graves and N Jaitly, · 2014
Earlier work this paper cites.
“On the properties of neural machine translation: Encoder-decoder approaches,”
K Cho, Bart Van M, D Bahdanau, and Y Bengio, · 2014
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
N Srivastava, G Hinton, A Krizhevsky, I Sutskever, and R Salakhutdinov, · 2014
Earlier work this paper cites.
“First-pass large vocabulary continuous speech recognition using bi-directional recurrent dnns,”
A Hannun, A Maas, D Jurafsky, and A Ng, · 2014
Earlier work this paper cites.
“The ibm 2015 english conversational telephone speech recognition system,”
G Saon, HK J Kuo, S Rennie, and M Picheny, · 2015
Earlier work this paper cites.
“Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,”
Y Miao, M Gowayyed, and F Metze, · 2015
Cited alongside, same era.
“Sequence level training with recurrent neural networks,”
M Ranzato, S Chopra, M Auli, and W Zaremba, · 2015
Cited alongside, same era.
“Librispeech: an asr corpus based on public domain audio books,”
V Panayotov, G Chen, D Povey, and S Khudanpur, · 2015
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift.,”
S Ioffe and C Szegedy, · 2015
Cited alongside, same era.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
K He, X Zhang, S Ren, and J Sun, · 2015
Cited alongside, same era.
“Xception: Deep learning with depthwise separable convolutions,”
F Chollet, · 2016
Later among the works it cites.
“Deep residual learning for image recognition,”
K He, X Zhang, S Ren, and J Sun, · 2016
Later among the works it cites.
“A theoretically grounded application of dropout in recurrent neural networks,”
Y Gal and Z Ghahramani, · 2016
Later among the works it cites.
“On multiplicative integration with recurrent neural networks,”
Y Wu, S Zhang, Y Zhang, Y Bengio, and R Salakhutdinov, · 2016
Later among the works it cites.
“Towards better decoding and language model integration in sequence to sequence models,”
J Chorowski and N Jaitly, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al., · 2016
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
D Bahdanau, J Chorowski, D Serdyuk, P Brakel, and Y Bengio, · 2016
Cited alongside, same era.
“An actor-critic algorithm for sequence prediction,”
D Bahdanau, P Brakel, K Xu, A Goyal, R Lowe, J Pineau, A Courville, and Y Bengio, · 2016
Cited alongside, same era.
“Self-critical sequence training for image captioning,”
S J Rennie, E Marcheret, Y Mroueh, J Ross, and V Goel, · 2016
Cited alongside, same era.
R Collobert, C Puhrsch, and G Synnaeve, · 2016
Later among the works it cites.
“The microsoft 2016 conversational speech recognition system,”
W Xiong, J Droppo, X Huang, F Seide, M Seltzer, A Stolcke, D Yu, and G Zweig, · 2017
Closest in time.
“A deep reinforced model for abstractive summarization,”
R Paulus, C Xiong, and R Socher, · 2017
Closest in time.
“Learning online alignments with continuous rewards policy gradient,”
Y Luo, C Chiu, N Jaitly, and I Sutskever, · 2017
Closest in time.