Fetching the paper…
Reading the bibliography…
Neural transducers have been widely used in automatic speech recognition (ASR).
H. Ney, “Speech translation: Coupling of recognition and translation,” in Proceedings of ICASSP , 1999, pp. 517–520
1999
Earlier work this paper cites.
E. Matusov, S. Kanthak, and H. Ney, “On the integration of speech recognition and statistical machine translation,” in European Conference on Speech Communicaton and Technology , 2005
2005
Earlier work this paper cites.
2012
Earlier work this paper cites.
J. Li, D. Yu, J.-T. Huang, and Y. Gong, “Improving wideband speech recognition using mixed-bandwidth training data in CD-DNN-HMM,” in Proceedings of SLT . IEEE, 2012, pp. 131–136
2012
Earlier work this paper cites.
M. Post, G. Kumar, A. Lopez, D. Karakos, C. Callison-Burch, and S. Khudanpur, “Improved speech-to-text translation with the fisher and callhome spanish-english speech translation corpus,” in Proceedings of IWSLT , 2013
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Berard, O. Pietquin, C. Servan, and L. Besacier, “Listen and translate: A proof of concept for end-to-end speech-to-text translation,” in NIPS Workshop on End-to-end Learning for Speech and Audio Processing , 2016
2016
Earlier work this paper cites.
C. Federmann and W. D. Lewis, “Microsoft speech language translation (MSLT) corpus: The IWSLT 2016 release for english, french and german,” in Proceedings of IWSLT , 2016
2016
Earlier work this paper cites.
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, “Sequence-to-sequence models can directly translate foreign speech,” Proceedings of Interspeech , pp. 2625–2629, 2017
2017
Earlier work this paper cites.
C. Raffel, M.-T. Luong, P. J. Liu, R. J. Weiss, and D. Eck, “Online and linear-time attention by enforcing monotonic alignments,” in Proceedings of the International Conference on Machine Learning , 2017, pp. 2837–2846
2017
Earlier work this paper cites.
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, “A comparison of sequence-to-sequence models for speech recognition,” in Proceedings of Interspeech , 2017, pp. 939–943
2017
Earlier work this paper cites.
L. C. Vila, C. Escolano, J. A. Fonollosa, and M. R. Costa-Jussa, “End-to-end speech translation with the transformer.” in Proceedings of Interspeech , 2018, pp. 60–63
2018
Earlier work this paper cites.
A. Bérard, L. Besacier, A. C. Kocabiyikoglu, and O. Pietquin, “End-to-end automatic speech translation of audiobooks,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2018, pp. 6224–6228
2018
Cited alongside, same era.
C.-C. Chiu and C. Raffel, “Monotonic chunkwise attention,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
N. Arivazhagan, C. Cherry, W. Macherey, C.-C. Chiu, S. Yavuz, R. Pang, W. Li, and C. Raffel, “Monotonic infinite lookback attention for simultaneous machine translation,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics , 2019, pp. 1313–1323
2019
Cited alongside, same era.
X. Ma, J. M. Pino, J. Cross, L. Puzon, and J. Gu, “Monotonic multihead attention,” in Proceedings of International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
J. Li, R. Zhao, Z. Meng, Y. Liu, W. Wei, S. Parthasarathy, V. Mazalov, Z. Wang, L. He, S. Zhao, and et al, “Developing rnnt models surpassing high-performance hybrid models with customization capability,” in Proceedings of Interspeech , 2020, pp. 3590–3594
2020
Later among the works it cites.
M. Gaido, M. A. Di Gangi, M. Negri, and M. Turchi, “End-to-end speech-translation with knowledge distillation,” in Proceedings of the International Conference on Spoken Language Translation , 2020, pp. 80–88
2020
Later among the works it cites.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2020, pp. 7829–7833
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Miao, G. Cheng, P. Zhang, T. Li, and Y. Yan, “Online hybrid ctc/attention architecture for end-to-end speech recognition,” Proceedings of Interspeech , pp. 2623–2627, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Jia, M. Johnson, W. Macherey, R. J. Weiss, Y. Cao, C.-C. Chiu, N. Ari, S. Laurenzo, and Y. Wu, “Leveraging weakly supervised data to improve end-to-end speech-to-text translation,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2019, pp. 7180–7184
2019
Cited alongside, same era.
M. Ma, L. Huang, H. Xiong, R. Zheng, K. Liu, B. Zheng, C. Zhang, Z. He, H. Liu, X. Li et al. , “Stacl: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics , 2019, pp. 3025–3036
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. Sperber and M. Paulik, “Speech translation and the end-to-end promise: Taking stock of where we are,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 7409–7421
2020
Cited alongside, same era.
H. Inaguma, Y. Gaur, L. Lu, J. Li, and Y. Gong, “Minimum latency training strategies for streaming sequence-to-sequence asr,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2020, pp. 6064–6068
2020
Cited alongside, same era.
T. N. Sainath, Y. He, B. Li, A. Narayanan, R. Pang, A. Bruguier, S.-Y. Chang, W. Li, R. Alvarez, Z. Chen, and et al, “A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,” in Proceedings of ICASSP , 2020, pp. 6059–6003
2020
Cited alongside, same era.
2020
Later among the works it cites.
P. Wang, T. N. Sainath, and R. J. Weiss, “Multitask training with text data for end-to-end speech recognition,” in Proc. of Interspeech , 2021, pp. 2566–2570
2021
Later among the works it cites.
X. Ma, Y. Wang, M. J. Dousti, P. Koehn, and J. Pino, “Streaming simultaneous speech translation with augmented memory transformer,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2021, pp. 7523–7527
2021
Later among the works it cites.
G. Saon, Z. Tüske, D. Bolanos, and B. Kingsbury, “Advancing rnn transducer technology for speech recognition,” in Proceedings of ICASSP , 2021, pp. 5654–5658
2021
Later among the works it cites.
2021
Later among the works it cites.
D. Liu, M. Du, X. Li, Y. Li, and E. Chen, “Cross attention augmented transducer networks for simultaneous translation,” in Proceedings of EMNLP , 2021, pp. 39–55
2021
Later among the works it cites.
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, “Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2021, pp. 5904–5908
2021
Later among the works it cites.