Fetching the paper…
Reading the bibliography…
End-to-end Automatic Speech Recognition (ASR) models are usually trained to optimize the loss of the whole token sequence, while neglecting explicit phonemic-granularity supervision.
A. Graves, S. Fernandez, F. Gomez, et al., “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML
2006
Earlier work this paper cites.
K. Gorman, J. Howell, and M. Wagner, “Prosodylab-aligner: A tool for forced alignment of laboratory speech,” Canadian Acoustics
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, et al., “The kaldi speech recognition toolkit,” in Proc. ASRU
2011
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, et al., “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. ICASSP
2016
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, et al., “End-to-end attention-based large vocabulary speech recognition,” in Proc. ICASSP
2016
Earlier work this paper cites.
S. Amodei, R. Ananthanarayanan, J. Anubhai, et al., “Deep speech 2: End-to-end speech recognition in English and Mandarin,” in Proc. ICML
2016
Earlier work this paper cites.
R. Prabhavalkar, K. Rao, T. Sainath, et al., “A comparison of sequence-to-sequence models for speech recognition,” in Proc. Interspeech
2017
Earlier work this paper cites.
Y. Belinkov, and J. Glass, “Analyzing hidden representations in end-to-end automatic speech recognition systems,” in Proc. NeurIPS
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, et al., “AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline,” in Proc. O-COCOSDA
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
C. Wang, Y. Wu, Y. Du, et al., “Semantic mask for transformer based end-to-end speech recognition,” in Proc. Interspeech
2019
Earlier work this paper cites.
Y. Chung, W. Hsu, H. Tang, et al., “An unsupervised autoregressive model for speech representation learning,” in Proc. Interspeech
2019
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, et al., “Wav2vec: Unsupervised pre-training for speech recognition,” in Proc. Interspeech
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Huang, Y. Lu, L. Wang, et al., “Exploring model units and training strategies for end-to-end speech recognition,” in Proc. ASRU
2019
Cited alongside, same era.
J. Li, Y. Wu, Y. Gaur, et al., “On the comparison of popular end-to-end models for large scale speech recognition,” in Proc. Interspeech
2020
Cited alongside, same era.
A. Fang, S. Filice, N. Limsopatham, et al., “Using phoneme representations to build predictive models robust to ASR errors,” in Proc. SIGIR
2020
Cited alongside, same era.
A. Baevski, and A. Mohamed, “Effectiveness of self-supervised pre-training for ASR,” in Proc. ICASSP
2020
Cited alongside, same era.
M. Ravanelli, J. Zhong, S. Pascual, et al., “Multi-task self-supervised learning for robust speech recognition,” in Proc. ICASSP
2020
Cited alongside, same era.
W. Hsu, B. Bolte, Y. Tsai, et al., “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. Audio, Speech, Language Process
2021
Closest in time.
A. Pasad, J. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in Proc. ASRU
2021
Closest in time.
C. Wang, Y. Wu, Y. Qian, et al., “UniSpeech: Unified speech representation learning with labeled and unlabeled data,” in Proc. ICML
2021
Closest in time.
2021
Closest in time.
L. Fu, X. Li, L. Zi, et al., “Incremental learning for end-to-end automatic speech recognition,” in Proc. ASRU
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Liu, S. Yang, P. Chi, et al., “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,” in Proc. ICASSP
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, et al., “Wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS
2020
Cited alongside, same era.
P. Khosla, P. Teterwak, C. Wang, et al., “Supervised contrastive learning,” in Proc. NeurIPS
2020
Cited alongside, same era.
A. Baevski, S. Schneider, and M. Auli, “Vq-wav2vec: Self-supervised learning of discrete speech representations,” in Proc. ICLR
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, et al., “A simple framework for contrastive learning of visual representations,” in Proc. ICML
2020
Cited alongside, same era.
Y. Zhang, J. Qin, D. Park, et al., “Pushing the limits of semi-supervised learning for automatic speech recognition,” in Proc. NeurIPS
2020
Cited alongside, same era.
J. Li, “Recent advances in end-to-end automatic speech recognition,” arXiv preprint arXiv:2111.01690
2021
Cited alongside, same era.
2021
Closest in time.
C. Talnikar, T. Likhomanenko, R. Collobert, et al., “Joint masked CPC and CTC training for ASR,” in Proc. ICASSP
2021
Closest in time.
2021
Closest in time.
Z. Yao, D. Wu, X. Wang, et al., “WeNet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,” in Proc. Interspeech
2021
Closest in time.
2021
Closest in time.
2022
Closest in time.
J. Bai, B. Li, Y. Zhang, et al., “Joint unsupervised and supervised training for multilingual asr,” in Proc. ICASSP
2022
Closest in time.