Fetching the paper…
Reading the bibliography…
Speech recognition technologies are gaining enormous popularity in various industrial applications.
“Sources of degradation of speech recognition in the telephone network,”
M. Pedro and S. Richard, · 1994
Earlier work this paper cites.
“Effect of speaking style on lvcsr performance,”
W. Mitch, T. Kelsey, H. Kate, and S. Amy, · 1996
Earlier work this paper cites.
“Towards lower error rates in phoneme recognition,”
S. Petr, M. Pavel, and Č. Jan, · 2004
Earlier work this paper cites.
“HKUST/MTS: A very large scale mandarin telephone speech corpus,”
Y. Liu, P. Fung, Y. Yang, C. Cieri, and et al, · 2006
Earlier work this paper cites.
“Automatic speech recognition and speech variability: A review,”
B. Mohamed, D. Renato, D. Olivier, D. Stephane, and et al, · 2007
Earlier work this paper cites.
“Unsupervised visual representation learning by context prediction,”
C. Doersch, A. Gupta, and A. Efros, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
“Purely sequence-trained neural networks for ASR based on lattice-free MMI,”
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, and et al, · 2016
Earlier work this paper cites.
“Very deep convolutional neural networks for noise robust speech recognition,”
Y. Qian, M. Bi, T. Tan, and K. Yu, · 2016
Earlier work this paper cites.
“AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline,”
H. Bu, J. Du, X. Na, B. Wu, and et al, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, and et al, · 2017
Earlier work this paper cites.
“Joint ctc-attention based end-to-end speech recognition using multi-task learning,”
K. Suyoun, H. Takaaki, and W. Shinji, · 2017
Cited alongside, same era.
“Improving language understanding with unsupervised learning,”
R. Alec, N. Karthik, S. Tim, and S. Ilya, · 2018
Cited alongside, same era.
Z. Lian, Y. Li, J.Tao, and J. Huang, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
O. Aaron van den, Y. Li, and V. Oriol, · 2018
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
Cited alongside, same era.
“Pre-trained language model representations for language generation,”
S. Edunov, A. Baevski, and M. Auli, · 2019
Closest in time.
“wav2vec: Unsupervised pre-training for speech recognition,”
S. Steffen, B. Alexei, C. Ronan, and A. Michael, · 2019
Closest in time.
“Learning speaker representations with mutual information,”
R. Mirco and B. Yoshua, · 2019
Closest in time.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
P. Santiago, R. Mirco, S. Joan, B. Antonio, and et al, · 2019
Closest in time.
“An unsupervised autoregressive model for speech representation learning,”
C. Yu-An, H. Wei-Ning, T. Hao, and G. James, · 2019
Closest in time.
“Self-attention aligner: A latency-control end-to-end model for ASR using self-attention network and chunk-hopping,”
L. Dong, F. Wang, and B. Xu, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“A comparison of modeling units in sequence-to-sequence speech recognition with the transformer on mandarin chinese,”
S. Zhou, L. Dong, S. Xu, and B. Xu, · 2018
Cited alongside, same era.
“Deep contextualized word representations,”
M. Peters, M. Neumann, M. Iyyer, M. Gardner, and et al, · 2018
Cited alongside, same era.
“Extending recurrent neural aligner for streaming end-to-end speech recognition in mandarin,”
L. Dong, S. Zhou, W. Chen, and B. Xu, · 2018
Cited alongside, same era.
“Primewords Chinese Corpus Set 1,” 2018,
Primewords Information Technology Co., Ltd., · 2018
Cited alongside, same era.
“Comparable study of modeling units for end-to-end mandarin speech recognition,”
W. Zou, D. Jiang, S. Zhao, G. Yang, and et al, · 2018
Cited alongside, same era.
“Multi-dialect speech recognition with a single sequence-to-sequence model,”
L. Bo, S. Tara N, S. Khe Chai, B. Michiel, and et al, · 2018
Cited alongside, same era.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M. Chang, K. Lee, and K. Toutanova, · 2019
Cited alongside, same era.
Closest in time.
“A comparative study on transformer vs rnn in speech applications,”
K. Shigeki, C. Nanxin, H. Tomoki, H. Takaaki, and et al, · 2019
Closest in time.
“Language models are unsupervised multitask learners,”
R. Alec, J. Wu, C. Rewon, D. Luan, and et al, · 2019
Closest in time.
“Roberta: A robustly optimized bert pretraining approach,”
Y. Liu, O. Myle, G. Naman, J. Du, and et al, · 2019
Closest in time.
“MAGICDATA Mandarin Chinese Read Speech Corpus,” http://www.imagicdatatech.com/index.php/home/dataopensource/data_info/id/101 , 2019
Magic Data Technology Co., Ltd, · 2019
Closest in time.
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, and et al, · 2019
Closest in time.
“Xlnet: Generalized autoregressive pretraining for language understanding,” 2019
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, and et al, · 2019
Closest in time.