Fetching the paper…
Reading the bibliography…
Non-autoregressive (NAR) models have achieved a large inference computation reduction and comparable results with autoregressive (AR) models on various sequence to sequence tasks.
E. A. Chi, J. Salazar, and K. Kirchhoff, “Align-refine: Non-autoregressive speech recognition via iterative realignment,” in Proc. NAACL . ACL, 2021, pp. 1920–1927
1927
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML . PMLR, 2006, pp. 369–376
2006
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. ICASSP . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell,” in Proc. ICASSP . IEEE, 2016, pp. 4960–4964
2016
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. ICASSP . IEEE, 2016, pp. 31–35
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones et al. , “Attention is all you need,” in Proc. NeurIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. ICASSP . IEEE, 2017, pp. 4835–4839
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proc. ICASSP . IEEE, 2017, pp. 241–245
2017
Earlier work this paper cites.
L. Dong, S. Xu, and B. Xu, “Speech-Transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in Proc. ICASSP . IEEE, 2018, pp. 5884–5888
2018
Earlier work this paper cites.
J. Gu, J. Bradbury, C. Xiong, V. O. Li, and R. Socher, “Non-autoregressive neural machine translation,” in Proc. ICLR , 2018
2018
Earlier work this paper cites.
J. Libovickỳ and J. Helcl, “End-to-end non-autoregressive neural machine translation with connectionist temporal classification,” in Proc. EMNLP . ACL, 2018, pp. 3016–3021
2018
Earlier work this paper cites.
J. Lee, E. Mansimov, and K. Cho, “Deterministic non-autoregressive neural sequence modeling by iterative refinement,” in Proc. EMNLP . ACL, 2018, pp. 1173–1182
2018
Earlier work this paper cites.
Y. Qian, X. Chang, and D. Yu, “Single-channel multi-talker speech recognition with permutation invariant training,” Speech Communication , vol. 104, pp. 1–11, 2018
2018
Earlier work this paper cites.
H. Seki, T. Hori, S. Watanabe, J. L. Roux, and J. R. Hershey, “A purely end-to-end system for multi-speaker speech recognition,” in Proc. ACL , 2018, pp. 2620–2630
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen et al. , “ESPnet: End-to-end speech processing toolkit,” in Proc. Interspeech . ISCA, 2018, pp. 2207–2211
2018
Cited alongside, same era.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma et al. , “A comparative study on Transformer vs RNN in speech applications,” in Proc. ASRU . IEEE, 2019, pp. 449–456
2019
Cited alongside, same era.
M. Stern, W. Chan, J. Kiros, and J. Uszkoreit, “Insertion Transformer: Flexible sequence generation via insertion operations,” in Proc. ICML . PMLR, 2019, pp. 5976–5985
2019
Cited alongside, same era.
Y. Higuchi, S. Watanabe, N. Chen, T. Ogawa, and T. Kobayashi, “Mask CTC: Non-autoregressive end-to-end ASR with CTC and mask predict,” in Proc. Interspeech . ISCA, 2020, pp. 3655–3659
2020
Later among the works it cites.
W. Chan, C. Saharia, G. Hinton, M. Norouzi, and N. Jaitly, “Imputer: Sequence modelling via imputation and dynamic programming,” in Proc. ICML . PMLR, 2020, pp. 1403–1413
2020
Later among the works it cites.
Z. Tian, J. Yi, J. Tao, Y. Bai, S. Zhang, and Z. Wen, “Spike-triggered non-autoregressive Transformer for end-to-end speech recognition,” in Proc. Interspeech . ISCA, 2020, pp. 5026–5020
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
M. Ghazvininejad, O. Levy, Y. Liu, and L. Zettlemoyer, “Mask-predict: Parallel decoding of conditional masked language models,” in Proc. EMNLP-IJCNLP . ACL, 2019, pp. 6112–6121
2019
Cited alongside, same era.
2019
Cited alongside, same era.
T. Menne, I. Sklyar, R. Schlüter, and H. Ney, “Analysis of deep clustering as preprocessing for automatic speech recognition of sparsely overlapping speech.” ISCA, 2019, pp. 2638–2642
2019
Cited alongside, same era.
X. Chang, Y. Qian, K. Yu, and S. Watanabe, “End-to-end monaural multi-speaker asr system without pretraining,” in Proc. ICASSP . IEEE, 2019, pp. 6256–6260
2019
Cited alongside, same era.
T. von Neumann, K. Kinoshita, M. Delcroix, S. Araki, T. Nakatani, and R. Haeb-Umbach, “All-neural online source separation, counting, and diarization for meeting analysis,” in Proc. ICASSP . IEEE, 2019, pp. 91–95
2019
Cited alongside, same era.
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. L. Roux, “Wham!: Extending speech separation to noisy environments,” in Proc. Interspeech . ISCA, 2019, pp. 1368–1372
2019
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech . ISCA, 2020, pp. 5036–5040
2020
Cited alongside, same era.
X. Chang, W. Zhang, Y. Qian, J. Le Roux, and S. Watanabe, “End-to-end multi-speaker speech recognition with Transformer,” in Proc. ICASSP . IEEE, 2020, pp. 6134–6138
2020
Later among the works it cites.
N. Kanda, Y. Gaur, X. Wang, Z. Meng, Z. Chen, T. Zhou, and T. Yoshioka, “Joint speaker counting, speech recognition, and speaker identification for overlapped speech of any number of speakers,” in Proc. Interspeech . ISCA, 2020, pp. 36–40
2020
Later among the works it cites.
T. von Neumann, K. Kinoshita, L. Drude, C. Boeddeker, M. Delcroix, T. Nakatani, and R. Haeb-Umbach, “End-to-end training of time domain audio separation and recognition,” in Proc. ICASSP . IEEE, 2020, pp. 7004–7008
2020
Later among the works it cites.
J. Shi, J. Xu, Y. Fujita, S. Watanabe, and B. Xu, “Speaker-conditional chain model for speech separation and extraction,” in Proc. Interspeech . ISCA, 2020, pp. 2707–2711
2020
Later among the works it cites.
J. Shi, X. Chang, P. Guo, S. Watanabe, Y. Fujita, J. Xu, B. Xu, and L. Xie, “Sequence to multi-sequence learning via conditional chain mapping for mixture signals,” in Proc. NeurIPS , 2020, pp. 3735–3747
2020
Later among the works it cites.
2020
Later among the works it cites.
P. Guo, F. Boyer, X. Chang, T. Hayashi, Y. Higuchi, H. Inaguma, N. Kamo, C. Li, D. Garcia-Romero, J. Shi et al. , “Recent developments on ESPnet toolkit boosted by Conformer,” in Proc. ICASSP . IEEE, 2021, pp. 5874–5878
2021
Closest in time.
Y. Higuchi, H. Inaguma, S. Watanabe, T. Ogawa, and T. Kobayashi, “Improved Mask-CTC for non-autoregressive end-to-end ASR,” in Proc. ICASSP . IEEE, 2021, pp. 8363–8367
2021
Closest in time.
J. Lee and S. Watanabe, “Intermediate loss regularization for ctc-based speech recognition,” in Proc. ICASSP . IEEE, 2021, pp. 6224–6228
2021
Closest in time.