Fetching the paper…
Reading the bibliography…
This paper proposes a method to relax the conditional independence assumption of connectionist temporal classification (CTC)-based automatic speech recognition (ASR) models.
E. A. Chi, J. Salazar, and K. Kirchhoff, “Align-refine: Non-autoregressive speech recognition via iterative realignment,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 1920–1927
1927
Earlier work this paper cites.
D. B. Paul and J. Baker, “The design for the Wall Street Journal-based CSR corpus,” in Proceedings of Workshop on Speech and Natural Language , 1992
1992
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of International Conference on Machine Learning (ICML) . PMLR, 2006, pp. 369–376
2006
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in Proceedings of International Conference on Machine Learning (ICML) . PMLR, 2014, pp. 1764–1772
2014
Earlier work this paper cites.
A. Rousseau, P. Deléglise, Y. Esteve et al. , “Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks.” in Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC) , 2014, pp. 3935–3939
2014
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” Proc. Interspeech , 2015
2015
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 4945–4949
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” stat , vol. 1050, p. 21, 2016
2016
Earlier work this paper cites.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in European conference on computer vision (ECCV) , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in 2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment (O-COCOSDA) . IEEE, 2017, pp. 1–5
2017
Earlier work this paper cites.
C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina et al. , “State-of-the-art speech recognition with sequence-to-sequence models,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4774–4778
2018
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5884–5888
2018
Cited alongside, same era.
J. Gu, J. Bradbury, C. Xiong, V. O. Li, and R. Socher, “Non-autoregressive neural machine translation,” Proceedings of International Conference on Learning Representations (ICLR) , 2018
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N.-E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen et al. , “Espnet: End-to-end speech processing toolkit,” Proc. Interspeech 2018 , pp. 2207–2211, 2018
S. Kriman, S. Beliaev, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, and Y. Zhang, “Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions,” in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6124–6128
2020
Later among the works it cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” Proc. Interspeech 2020 , pp. 5036–5040, 2020
2020
Later among the works it cites.
Y. Higuchi, S. Watanabe, N. Chen, T. Ogawa, and T. Kobayashi, “Mask CTC: non-autoregressive end-to-end ASR with CTC and Mask Predict,” Proc. Interspeech 2020 , pp. 3655–3659, 2020
2020
Later among the works it cites.
Y. Fujita, S. Watanabe, M. Omachi, and X. Chang, “Insertion-based modeling for end-to-end automatic speech recognition,” Proc. Interspeech 2020 , pp. 3660–3664, 2020
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
T. Kudo and J. Richardson, “Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2018, pp. 66–71
2018
Cited alongside, same era.
M. Ott, S. Edunov, D. Grangier, and M. Auli, “Scaling neural machine translation,” in Proceedings of the Third Conference on Machine Translation: Research Papers , 2018, pp. 1–9
2018
Cited alongside, same era.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang et al. , “A comparative study on transformer vs rnn in speech applications,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 449–456
2019
Cited alongside, same era.
S. Karita, N. E. Y. Soplin, S. Watanabe, M. Delcroix, A. Ogawa, and T. Nakatani, “Improving transformer-based end-to-end speech recognition with connectionist temporal classification and language model integration,” Proc. Interspeech 2019 , pp. 1408–1412, 2019
2019
Cited alongside, same era.
M. Ghazvininejad, O. Levy, Y. Liu, and L. Zettlemoyer, “Mask-predict: Parallel decoding of conditional masked language models,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 6114–6123
2019
Cited alongside, same era.
M. Stern, W. Chan, J. Kiros, and J. Uszkoreit, “Insertion transformer: Flexible sequence generation via insertion operations,” in Proceedings of International Conference on Machine Learning (ICML) . PMLR, 2019, pp. 5976–5985
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” Proc. Interspeech 2019 , pp. 2613–2617, 2019
2019
Cited alongside, same era.
Later among the works it cites.
W. Chan, C. Saharia, G. Hinton, M. Norouzi, and N. Jaitly, “Imputer: Sequence modelling via imputation and dynamic programming,” in Proceedings of International Conference on Machine Learning (ICML) . PMLR, 2020, pp. 1403–1413
2020
Later among the works it cites.
Y. Bai, J. Yi, J. Tao, Z. Tian, Z. Wen, and S. Zhang, “Listen attentively, and spell once: Whole sentence generation via a non-autoregressive architecture for low-latency speech recognition,” Proc. Interspeech 2020 , pp. 3381–3385, 2020
2020
Later among the works it cites.
Z. Tian, J. Yi, J. Tao, Y. Bai, S. Zhang, and Z. Wen, “Spike-triggered non-autoregressive transformer for end-to-end speech recognition,” Proc. Interspeech 2020 , pp. 5026–5030, 2020
2020
Later among the works it cites.
N. Chen, S. Watanabe, J. Villalba, and N. Dehak, “Listen and fill in the missing letters: Non-autoregressive transformer for speech recognition,” IEEE Signal Processing Letters , vol. 28, pp. 121–125, 2021
2021
Closest in time.
Y. Higuchi, H. Inaguma, S. Watanabe, T. Ogawa, and T. Kobayashi, “Improved Mask-CTC for non-autoregressive end-to-end ASR,” in 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021
2021
Closest in time.
J. Lee and S. Watanabe, “Intermediate loss regularization for CTC-based speech recognition,” in 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021
2021
Closest in time.
R. Fan, W. Chu, P. Chang, and J. Xiao, “CASS-NAT: CTC alignment-based single step non-autoregressive transformer for speech recognition,” in 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021
2021
Closest in time.