Fetching the paper…
Reading the bibliography…
End-to-end models reach state-of-the-art performance for speech recognition, but global soft attention is not monotonic, which might lead to convergence problems, to instability, to bad generalisation, cannot be used for online streaming, and is also inefficient in calculation.
A continuous speech recognition system embedding mlp into hmm
Bourlard, H. and Morgan, N · 1990
Earlier work this paper cites.
Switchboard: Telephone speech corpus for research and development
Godfrey, J. J., Holliman, E. C., and McDaniel, J · 1992
Earlier work this paper cites.
An application of recurrent nets to phone probability estimation
Robinson, T · 1994
Earlier work this paper cites.
From HMM’s to segment models: A unified view of stochastic modeling for speech recognition
Ostendorf, M., Digalakis, V., and Kimball, O · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
A segmental CRF approach to large vocabulary continuous speech recognition
Zweig, G. and Nguyen, P · 2009
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Bengio, Y · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Graves, A · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A · 2013
Earlier work this paper cites.
Multiple object recognition with visual attention
Ba, J., Mnih, V., and Kavukcuoglu, K · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Mnih, V., Heess, N., Graves, A., et al · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Automatic Speech Recognition - A Deep Learning Approach
Yu, D. and Deng, L · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Pointer networks
Vinyals, O., Fortunato, M., and Jaitly, N · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y · 2015
Earlier work this paper cites.
Alignment-based neural machine translation
Alkhouli, T., Bretschner, G., Peter, J.-T., Hethnawi, M., Guta, A., and Ney, H · 2016
Earlier work this paper cites.
An online sequence-to-sequence model using partial conditioning
Jaitly, N., Le, Q. V., Vinyals, O., Sutskever, I., Sussillo, D., and Bengio, S · 2016
Earlier work this paper cites.
Segmental recurrent neural networks for end-to-end speech recognition
Lu, L., Kong, L., Dyer, C., Smith, N. A., and Renals, S · 2016
Cited alongside, same era.
Modeling coverage for neural machine translation
Tu, Z., Lu, Z., Liu, Y., Liu, X., and Li, H · 2016
Cited alongside, same era.
Towards online-recognition with deep bidirectional LSTM acoustic models
Zeyer, A., Schlüter, R., and Ney, H · 2016
Cited alongside, same era.
Biasing attention-based recurrent neural networks using external alignment information
Alkhouli, T. and Ney, H · 2017
Cited alongside, same era.
Variational attention for sequence-to-sequence models
Bahuleyan, H., Mou, L., Vechtomova, O., and Poupart, P · 2017
Cited alongside, same era.
Extending recurrent neural aligner for streaming end-to-end speech recognition in mandarin
Dong, L., Zhou, S., Chen, W., and Xu, B · 2018
Later among the works it cites.
An online attention-based model for speech recognition
Fan, R., Zhou, P., Chen, W., Jia, J., and Liu, G · 2018
Later among the works it cites.
Learning hard alignments with variational inference
Lawson, D., Chiu, C.-C., Tucker, G., Raffel, C., Swersky, K., and Jaitly, N · 2018
Later among the works it cites.
Minimum word error rate training for attention-based sequence-to-sequence models
Prabhavalkar, R., Sainath, T. N., Wu, Y., Nguyen, P., Chen, Z., Chiu, C.-C., and Kannan, A · 2018
Later among the works it cites.
Surprisingly easy hard-attention for sequence to sequence learning
Shankar, S., Garg, S., and Sarawagi, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring neural transducers for end-to-end speech recognition
Battenberg, E., Chen, J., Child, R., Coates, A., Li, Y. G. Y., Liu, H., Satheesh, S., Sriram, A., and Zhu, Z · 2017
Cited alongside, same era.
Chiu, C.-C. and Raffel, C · 2017
Cited alongside, same era.
Joint CTC/attention decoding for end-to-end speech recognition
Hori, T., Watanabe, S., and Hershey, J. R · 2017
Cited alongside, same era.
Gaussian prediction based attention for online end-to-end speech recognition
Hou, J., Zhang, S., and Dai, L.-R · 2017
Cited alongside, same era.
Joint CTC-attention based end-to-end speech recognition using multi-task learning
Kim, S., Hori, T., and Watanabe, S · 2017
Cited alongside, same era.
Low latency acoustic modeling using temporal convolution and lstms
Peddinti, V., Wang, Y., Povey, D., and Khudanpur, S · 2017
Cited alongside, same era.
Online and linear-time attention by enforcing monotonic alignments
Raffel, C., Luong, M.-T., Liu, P. J., Weiss, R. J., and Eck, D · 2017
Cited alongside, same era.
Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention
Tachibana, H., Uenoyama, K., and Aihara, S · 2018
Later among the works it cites.
Neural hidden markov model for machine translation
Wang, W., Zhu, D., Alkhouli, T., Gan, Z., and Ney, H · 2018
Later among the works it cites.
Hard non-monotonic attention for character-level transduction
Wu, S., Shapiro, P., and Cotterell, R · 2018
Later among the works it cites.
Forward attention in sequence-to-sequence acoustic modeling for speech synthesis
Zhang, J.-X., Ling, Z.-H., and Dai, L.-R · 2018
Later among the works it cites.
Monotonic infinite lookback attention for simultaneous machine translation
Arivazhagan, N., Cherry, C., Macherey, W., Chiu, C.-C., Yavuz, S., Pang, R., Li, W., and Raffel, C · 2019
Later among the works it cites.
Two-pass end-to-end speech recognition
Chiu, C.-C., Rybach, D., McGraw, I., Visontai, M., Liang, Q., Prabhavalkar, R., Pang, R., Sainath, T., Strohman, T., Li, W., He, Y. R., and Wu, Y · 2019
Later among the works it cites.
Robust sequence-to-sequence acoustic modeling with stepwise monotonic attention for neural TTS
He, M., Deng, Y., and He, L · 2019
Later among the works it cites.
An analysis of local monotonic attention variants
Merboldt, A., Zeyer, A., Schlüter, R., and Ney, H · 2019
Later among the works it cites.
Online hybrid CTC/attention architecture for end-to-end speech recognition
Miao, H., Cheng, G., Zhang, P., Li, T., and Yan, Y · 2019
Later among the works it cites.
Triggered attention for end-to-end speech recognition
Moritz, N., Hori, T., and Le Roux, J · 2019
Later among the works it cites.
SpecAugment: A simple data augmentation method for automatic speech recognition
Park, D. S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E. D., and Le, Q. V · 2019
Later among the works it cites.
Posterior attention models for sequence to sequence learning
Shankar, S. and Sarawagi, S · 2019
Later among the works it cites.
A comparison of Transformer and LSTM encoder decoder models for ASR
Zeyer, A., Bahar, P., Irie, K., Schlüter, R., and Ney, H · 2019
Later among the works it cites.
Windowed attention mechanisms for speech recognition
Zhang, S., Loweimi, E., Bell, P., and Renals, S · 2019
Later among the works it cites.
Tüske, Z., Saon, G., Audhkhasi, K., and Kingsbury, B · 2020
Later among the works it cites.
A new training pipeline for an improved neural transducer
Zeyer, A., Merboldt, A., Schlüter, R., and Ney, H · 2020
Later among the works it cites.