Fetching the paper…
Reading the bibliography…
In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model for large scale speech recognition applications.
“Subphonetic modeling with markov states-senone,”
M.-Y. Hwang and X. Huang, · 1992
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Lattice-based optimization of sequence classification criteria for neural-network acoustic modeling,”
B. Kingsbury, · 2009
Earlier work this paper cites.
“Emformer: efficient low latency streaming transformer acoustic model for speech recognition,”
Y. Shi, Y. Wang, C. Wu, C.-C. Yeh, et al., · 2010
Earlier work this paper cites.
“Conversational speech transcription using context-dependent deep neural networks,”
F. Seide, G. Li, and D. Yu, · 2011
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, et al., · 2011
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition,”
G. Hinton, L. Deng, D. Yu, et al., · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Japanese and korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling,”
H. Sak, A. Senior, and F. Beaufays, · 2014
Earlier work this paper cites.
“Convolutional neural networks for speech recognition,”
O. Abdel-Hamid, A. Mohamed, H. Jiang, et al., · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2014
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Cited alongside, same era.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
V. Peddinti, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“Feedforward sequential memory neural networks without recurrent feedback,”
S. Zhang, H. Jiang, S. Wei, and L. Dai, · 2015
Cited alongside, same era.
“Audio augmentation for speech recognition,”
T. Ko, V. Peddinti, D. Povey, et al., · 2015
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, et al., · 2019
Later among the works it cites.
“Transformer-based acoustic modeling for hybrid speech recognition,”
Y. Wang, A. Mohamed, D. Le, C. Liu, A. Xiao, et al., · 2020
Closest in time.
“Conformer: Convolution-augmented Transformer for Speech Recognition,”
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, et al., · 2020
Closest in time.
“Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,”
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, · 2020
Closest in time.
“Streaming automatic speech recognition with the transformer model,”
N. Moritz, T. Hori, and J. L. Roux, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Highway long short-term memory RNNs for distant speech recognition,”
Y. Zhang, G. Chen, D. Yu, et al., · 2016
Cited alongside, same era.
J. Lei Ba, J. Kiros, and G. E. Hinton, · 2016
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, et al., · 2017
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M.-W. Chang, K. Lee, et al., · 2018
Cited alongside, same era.
“A Comparative Study on Transformer vs RNN in Speech Applications,”
S. Karita, N. Chen, T. Hayashi, et al., · 2019
Cited alongside, same era.
“From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition,”
D. Le, X. Zhang, W. Zheng, et al., · 2019
Cited alongside, same era.
Connectionist speech recognition: a hybrid approach
H. A. Bourlard and N. Morgan,
Cited in the paper.
Closest in time.
“Streaming Transformer-based Acoustic Models Using Self-attention with Augmented Memory,”
C. Wu, Y. Wang, Y. Shi, C.-F. Yeh, and F. Zhang, · 2020
Closest in time.
“Deja-vu: Double Feature Presentation and Iterated loss in Deep Transformer Networks,”
A. Tjandra, C. Liu, F. Zhang, et al., · 2020
Closest in time.
“Benchmarking LF-MMI, CTC and RNN-T criteria for streaming ASR,”
X. Zhang, F. Zhang, C. Liu, K. Schubert, J. Chan, et al., · 2021
Closest in time.
“Improving RNN transducer based ASR with auxiliary tasks,”
C. Liu, F. Zhang, D. Le, S. Kim, et al., · 2021
Closest in time.
“Streaming Attention-Based Models with Augmented Memory for End-to-End Speech Recognition,”
C.-F. Yeh, Y. Wang, S. Yang, C. Wu, F. Zhang, et al., · 2021
Closest in time.