Fetching the paper…
Reading the bibliography…
In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state-of-the-art results.
M. Mohri, F. Pereira, and M. Riley, “Weighted Finite-State Transducers in Speech Recognition,”
2002
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,” in
2006
Earlier work this paper cites.
G. Dahl, D. Yu, L. Deng, and A. Acero, “Large Vocabulary Continuous Speech Recognition With Context-Dependent DBN-HMMS,” in
2011
Earlier work this paper cites.
A. Mohamed, G. E. Dahl, and G. Hinton, “Acoustic Modeling using Deep Belief Networks,”
2011
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl
2012
Earlier work this paper cites.
A. Graves, “Sequence Transduction with Recurrent Neural Networks,” in
2012
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and Korean voice search,” in
2012
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in
2014
Earlier work this paper cites.
O. Abdel-Hamid, A. Mohamed, H. Jiang
2014
Earlier work this paper cites.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Andrew, H. Sak, F. de Chaumont Quitry, T. Sainath
2015
Earlier work this paper cites.
H. Sak, A. Senior, K. Rao, O. İrsoy
2015
Earlier work this paper cites.
H. Sak, A. Senior, K. Rao, and F. Beaufays, “Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
Y. Zhang, G. Chen, D. Yu
2016
Cited alongside, same era.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani
2016
Cited alongside, same era.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai
2016
Cited alongside, same era.
D. Le, X. Zhang, W. Zheng, C. Fügen
2019
Later among the works it cites.
A. Hannun, A. Lee, Q. Xu, and R. Collobert, “Sequence-to-Sequence Speech Recognition with Time-Depth Separable Convolutions,”
2019
Later among the works it cites.
C. Lüscher, E. Beck, K. Irie
2019
Later among the works it cites.
Y. Wang, A. Mohamed, D. Le, C. Liu
2019
Later among the works it cites.
S. Karita, N. Chen, T. Hayashi
2019
Later among the works it cites.
G. Synnaeve, Q. Xu, J. Kahn, T. Likhomanenko
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Pundak and T. Sainath, “Lower Frame Rate Neural Network Acoustic Models,” in
2016
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
J. Lei Ba, J. Kiros, and G. E. Hinton, “Layer Normalization,”
2016
Cited alongside, same era.
R. Sennrich, B. Haddow, and A. Birch, “Neural Machine Translation of Rare Words with Subword Units,” in
2016
Cited alongside, same era.
Y. Wu, M. Schuster, Z. Chen, Q. V. Le
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar
2017
Cited alongside, same era.
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang
2019
Later among the works it cites.
O. Myle, E. Sergey, B. Alexei, F. Angela
2019
Later among the works it cites.
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Language Modeling with Deep Transformers,” in
2019
Later among the works it cites.
K. J. Han, R. Prieto, K. Wu, and T. Ma, “State-of-the-Art Speech Recognition Using Multi-Stream Self-Attention With Dilated 1D Convolutions,” in
2019
Later among the works it cites.
T. N. Sainath, Y. He, B. Li, A. Narayanan
2020
Closest in time.
Q. Zhang, H. Lu, H. Sak, A. Tripathi
2020
Closest in time.
A. Tjandra, C. Liu, F. Zhang
2020
Closest in time.
D. S. Park, Y. Zhang, C.-C. Chiu, Y. Chen
2020
Closest in time.