Fetching the paper…
Reading the bibliography…
Building a good speech recognition system usually requires large amounts of transcribed data, which is expensive to collect.
W. Mitch, T. Kelsey, H. Kate, and S. Amy, “Effect of speaking style on lvcsr performance,” in
1996
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: a resource for the next generations of speech-to-text,” in
2004
Earlier work this paper cites.
B. Mohamed, D. Renato, D. Olivier, D. Stephane, and et al, “Automatic speech recognition and speech variability: A review,”
2007
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in
2014
Earlier work this paper cites.
C. Doersch, A. Gupta, and A. Efros, “Unsupervised visual representation learning by context prediction,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Earlier work this paper cites.
X. Shi, I. Padhi, and K. Knight, “Does string-based neural mt learn source syntax?” in
2016
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, and et al., “Overcoming catastrophic forgetting in neural networks,”
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, and et al, “Attention is all you need,” in
2017
Earlier work this paper cites.
K. Suyoun, H. Takaaki, and W. Shinji, “Joint ctc-attention based end-to-end speech recognition using multi-task learning,” in
2017
Earlier work this paper cites.
O. A. van den, Y. Li, and V. Oriol, “Representation learning with contrastive predictive coding,”
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Howard and S. Ruder, “Universal language model fine-tuning for text classification,” in
2018
Cited alongside, same era.
T. N. Sainath, C.-C. Chiu, R. Prabhavalkar, A. Kannan, and et al, “Improving the performance of online neural transducer models,” in
2018
Cited alongside, same era.
W. Zou, D. Jiang, S. Zhao, G. Yang, and et al, “Comparable study of modeling units for end-to-end mandarin speech recognition,” in
2018
Cited alongside, same era.
T. Zenkel, R. Sanabria, F. Metze, and A. H. Waibel, “Subword and crossword units for ctc acoustic models,” in
2018
Cited alongside, same era.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in
2019
Cited alongside, same era.
A. Chronopoulou, C. Baziotis, and A. Potamianos, “An embarrassingly simple approach for transfer learning from pretrained language models,”
2019
Later among the works it cites.
C. Sun, X. Qiu, Y. Xu, and X. Huang, “How to fine-tune bert for text classification?”
2019
Later among the works it cites.
2019
Later among the works it cites.
L. Dong, N. Yang, W. Wang, F. Wei, and et al, “Unified language model pre-training for natural language understanding and generation,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
R. Mirco and B. Yoshua, “Learning speaker representations with mutual information,”
2019
Cited alongside, same era.
S. Steffen, B. Alexei, C. Ronan, and A. Michael, “wav2vec: Unsupervised pre-training for speech recognition,”
2019
Cited alongside, same era.
P. Santiago, R. Mirco, S. Joan, B. Antonio, and et al, “Learning problem-agnostic speech representations from multiple self-supervised tasks,”
2019
Cited alongside, same era.
C. Yu-An, H. Wei-Ning, T. Hao, and G. James, “An unsupervised autoregressive model for speech representation learning,”
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Salazar, K. Kirchhoff, and Z. Huang, “Self-attention networks for connectionist temporal classification in speech recognition,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, and et al, “Specaugment: A simple data augmentation method for automatic speech recognition,” in
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.