Fetching the paper…
Reading the bibliography…
We propose a new end-to-end neural acoustic model for automatic speech recognition.
“The design for the wall street journal based csr corpus,”
D. B. Paul and J. M. Baker, · 1992
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Rigid-motion scattering for texture classification,”
L Sifre and S. Mallat, · 2014
Earlier work this paper cites.
“Learning visual representations at scale,”
V. Vanhoucke, · 2014
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
S. Ioffe and C. Szegedy, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Squeezenet: Alexnet-level accuracy with 50x fewer parameters and <1mb model,”
F.N. Iandola, M. W. Moskewicz, K. Ashraf, S. Han, W.J. Dally, and K. Keutzer, · 2016
Earlier work this paper cites.
J. Ba, J.R. Kiros, and G.E. Hinton, · 2016
Earlier work this paper cites.
“Instance normalization: The missing ingredient for fast stylization,”
D. Ulyanov, A. Vedaldi, and V. Lempitsky, · 2016
Earlier work this paper cites.
“Towards better decoding and language model integration in sequence to sequence models,”
J. Chorowski and N. Jaitly, · 2016
Cited alongside, same era.
“Mobilenets: Efficient convolutional neural networks for mobile vision applications,”
A.G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, · 2017
Cited alongside, same era.
“Xception: Deep learning with depthwise separable convolutions,”
F. Chollet, · 2017
Cited alongside, same era.
“Letter-based speech recognition with gated convnets,”
V. Liptchinsky, G. Synnaeve, and R. Collobert, · 2017
Cited alongside, same era.
“Improved regularization of convolutional neural networks with cutout,”
“End-to-end speech recognition from the raw waveform,”
N. Zeghidour, N. Usunier, G. Synnaeve, R. Collobert, and E. Dupoux, · 2018
Later among the works it cites.
“Nemo: a toolkit for building ai applications using neural modules,”
O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V. Lavrukhin, J. Cook, et al., · 2019
Closest in time.
“Efficientnet: Rethinking model scaling for convolutional neural networks,”
M. Tan and Q.V. Le, · 2019
Closest in time.
“Sequence-to-Sequence Speech Recognition with Time-Depth Separable Convolutions,”
A. Hannun, A. Lee, Q. Xu, and R. Collobert, · 2019
Closest in time.
K.J. Han, R. Prieto, K. Wu, and T. Ma, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Devries and G.W. Taylor, · 2017
Cited alongside, same era.
P. Micikevicius, S. Narang, J. Alben, et al., · 2017
Cited alongside, same era.
“Shufflenet: An extremely efficient convolutional neural network for mobile devices,”
X. Zhang, X. Zhou, M. Lin, and J. Sun, · 2018
Cited alongside, same era.
“Mobilenetv2: Inverted residuals and linear bottlenecks,”
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L. Chen, · 2018
Cited alongside, same era.
“Group normalization,”
Y. Wu and K. He, · 2018
Cited alongside, same era.
“Fully convolutional speech recognition,”
N. Zeghidour, Q. Xu, V. Liptchinsky, N. Usunier, G. Synnaeve, and R. Collobert, · 2018
Cited alongside, same era.
Closest in time.
“Jasper: An end-to-end convolutional neural acoustic model,”
J. Li, V. Lavrukhin, B. Ginsburg, R. Leary, O. Kuchaiev, J.M. Cohen, H. Nguyen, and R.T. Gadde, · 2019
Closest in time.
“Common voice,” https://voice.mozilla.org/en
Mozilla, · 2019
Closest in time.
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Z. Dai, Z. Yang, Y. Yang, J.G. Carbonell, Q.V. Le, and R. Salakhutdinov, · 2019
Closest in time.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E.D. Cubuk, and Q.V. Le, · 2019
Closest in time.
“Stochastic gradient methods with layer-wise adaptive moments for training of deep networks,”
B. Ginsburg, P. Castonguay, O. Hrinchuk, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, H. Nguyen, and Cohen J. M., · 2019
Closest in time.