Fetching the paper…
Reading the bibliography…
In our previous work we demonstrated that a single headed attention encoder-decoder model is able to reach state-of-the-art results in conversational speech recognition.
Y. Nesterov, “A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2}) ,” in Soviet Mathematics Doklady , vol. 27, no. 2, 1983, pp. 372–376
1983
Earlier work this paper cites.
H. Hermansky, “Perceptual linear predictive (PLP) analysis of speech,” Journal of the Acoustical Society of America , vol. 87, no. 4, pp. 1738–1752, 1990
1990
Earlier work this paper cites.
A. Krogh and J. A. Hertz, “A simple weight decay can improve generalization,” in NIPS , 1992, pp. 950–957
1992
Earlier work this paper cites.
H. A. Bourlard and N. Morgan, Connectionist Speech Recognition: A Hybrid Approach . Norwell, MA, USA: Kluwer Academic Publishers, 1993
1993
Earlier work this paper cites.
A. F. Murray and P. J. Edwards, “Enhanced MLP performance and fault tolerance resulting from synaptic weight noise during training,” IEEE Trans. on Neural Networks , vol. 5, no. 5, pp. 792–802, 1994
1994
Earlier work this paper cites.
C. Dugast, L. Devillers, and X. Aubert, “Combining TDNN and HMM in a hybrid system for improved continuous-speech recognition,” IEEE Transactions on Speech and Audio Processing , vol. 2, no. 1, pp. 217–223, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
J. Kittler et al. , “On combining classifiers,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 20, no. 3, pp. 226–239, 1998
1998
Earlier work this paper cites.
K. Kirchhoff and J. A. Bilmes, “Combination and joint training of acoustic classifiers for speech recognition,” in ASR2000, ISCA Tutorial and Research Workshop , 2000, pp. 17–23
2000
Earlier work this paper cites.
R. Schlüter et al. , “Gammatone features and feature combination for large vocabulary speech recognition,” in ICASSP , 2007, pp. 649–652
2007
Earlier work this paper cites.
Y. Bengio et al. , “Curriculum learning,” in ICML , 2009, pp. 41–48
2009
Earlier work this paper cites.
D. Povey et al. , “The Kaldi speech recognition toolkit,” in ASRU , 2011
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
L. Wan et al. , “Regularization of neural networks using DropConnect,” in ICML , vol. 28, no. 3, 2013, pp. 1058–1066
2013
Earlier work this paper cites.
N. Kanda, R. Takeda, and Y. Obuchi, “Elastic spectral distortion for low resource speech recognition with deep neural networks,” in ASRU , 2013, pp. 309–314
2013
Earlier work this paper cites.
G. Saon et al. , “Speaker adaptation of neural network acoustic models using i-vectors,” in ASRU , 2013, pp. 55–59
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in ICLR , 2015
2015
Earlier work this paper cites.
T. Ko et al. , “Audio augmentation for speech recognition,” in Interspeech , 2015, pp. 3586–3589
2015
Earlier work this paper cites.
S. Bengio et al. , “Scheduled sampling for sequence prediction with recurrent neural networks,” in NIPS , 2015, pp. 1171–1179
2015
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Cited alongside, same era.
V. Manohar, D. Povey, and S. Khudanpur, “Semi-supervised maximum mutual information training of deep neural network acoustic models,” in Interspeech , 2015, pp. 2630–2634
2015
Cited alongside, same era.
2015
Cited alongside, same era.
J. K. Chorowski et al. , “Attention-based models for speech recognition,” in NIPS , 2015, pp. 577–585
2015
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,” in ICASSP , 2018, pp. 5884–5888
2018
Later among the works it cites.
Z. Tüske, R. Schlüter, and H. Ney, “Investigation on LSTM recurrent n-gram language models for speech recognition,” in Interspeech , 2018, pp. 3358–3362
2018
Later among the works it cites.
W. Xiong et al. , “The Microsoft 2017 conversational speech recognition system,” in ICASSP , 2018, pp. 5934–5938
2018
Later among the works it cites.
2018
Later among the works it cites.
G. Saon et al. , “Sequence noise injected training for end-to-end speech recognition,” in ICASSP , 2019, pp. 6261–6265
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Szegedy et al. , “Rethinking the inception architecture for computer vision,” in CVPR , 2016, pp. 2818–2826
2016
Cited alongside, same era.
D. Amodei et al. , “Deep speech 2: End-to-end speech recognition in English and Mandarin,” in ICML , 2016, pp. 173–182
2016
Cited alongside, same era.
Z. Zhu, J. H. Engel, and A. Hannun, “Learning multiscale features directly from waveforms,” in Interspeech , 2016, pp. 1305–1309
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Z. Tu et al. , “Modeling coverage for neural machine translation,” in ACL , 2016, pp. 76–85
2016
Cited alongside, same era.
W. Chan et al. , “Listen, attend and spell: a neural network for large vocabulary conversational speech recognition,” in ICASSP , 2016, pp. 4960–4964
2016
Cited alongside, same era.
D. Krueger et al. , “Zoneout: regularizing RNNs by randomly preserving hidden activations,” in ICLR , 2017
2017
Cited alongside, same era.
Later among the works it cites.
D. S. Park et al. , “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Interspeech , 2019, pp. 2613–2617
2019
Later among the works it cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Later among the works it cites.
M. Kitza et al. , “Cumulative adaptation for BLSTM acoustic models,” in Interspeech , 2019, pp. 754–758
2019
Later among the works it cites.
R. Al-Rfou et al. , “Character-level language modeling with deeper self-attention,” in AAAI Conf. on Artificial Intelligence , 2019, pp. 3159–3166
2019
Later among the works it cites.
Z. Dai et al. , “Transformer-XL: Attentive language models beyond a fixed-length context,” in ACL , 2019, pp. 2978–2988
2019
Later among the works it cites.
K. Irie et al. , “Training language models for long-span cross-sentence evaluation,” in ASRU , 2019, pp. 419–426
2019
Later among the works it cites.
E. McDermott, H. Sak, and E. Variani, “A density ratio approach to language model fusion in end-to-end automatic speech recognition,” in ASRU , 2019, pp. 434–441
2019
Later among the works it cites.
H. K. J. Kuo et al. , “End-to-end spoken language understanding without full transcripts,” in Interspeech , 2020, pp. 906–910
2020
Later among the works it cites.
A. Gulati et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Interspeech , 2020, pp. 5036–5040
2020
Later among the works it cites.
Z. Tüske et al. , “Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard,” in Interspeech , 2020, pp. 551–555
2020
Later among the works it cites.
C. Kim, K. Kim, and S. R. Indurthi, “Small energy masking for improved neural network training for end-to-end speech recognition,” in ICASSP , 2020, pp. 7684–7688
2020
Later among the works it cites.
Q. Zhang et al. , “Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,” in Interspeech , 2020, pp. 7829–7833
2020
Later among the works it cites.
G. Saon et al. , “Advancing RNN transducer technology for speech recognition,” in ICASSP , 2021
2021
Closest in time.