Fetching the paper…
Reading the bibliography…
Recently, Transformer based end-to-end models have achieved great success in many areas including speech recognition.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Highway long short-term memory rnns for distant speech recognition,”
Yu Zhang, Guoguo Chen, Dong Yu, Kaisheng Yaco, Sanjeev Khudanpur, and James Glass, · 2016
Earlier work this paper cites.
“A comparison of sequence-to-sequence models for speech recognition,”
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, · 2017
Earlier work this paper cites.
“Exploring neural transducers for end-to-end speech recognition,”
Eric Battenberg, Jitong Chen, et al., · 2017
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Sainath, et al., · 2018
Earlier work this paper cites.
“Advancing acoustic-to-word CTC model,”
J. Li, G. Ye, A. Das, R. Zhao, and Y. Gong, · 2018
Cited alongside, same era.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2018
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Self-attention with relative position representations,”
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani, · 2018
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, , et al., · 2019
Cited alongside, same era.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Cited alongside, same era.
Yangyang Shi, Yongqiang Wang, Chunyang Wu, Ching-Feng Yeh, Julian Chan, Frank Zhang, Duc Le, and Mike Seltzer, · 2020
Closest in time.
“Reducing the latency of end-to-end streaming speech recognition models with a scout network,”
Chengyi Wang, Yu Wu, Shujie Liu, Jinyu Li, Liang Lu, Guoli Ye, and Ming Zhou, · 2020
Closest in time.
“On the comparison of popular end-to-end models for large scale speech recognition,”
Jinyu Li, Yu Wu, Yashesh Gaur, Chengyi Wang, Rui Zhao, and Shujie Liu, · 2020
Closest in time.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Qian Zhang, Han Lu, et al., · 2020
Closest in time.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, et al., · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mahaveer Jain, Kjell Schubert, Jay Mahadeokar, et al., · 2019
Cited alongside, same era.
“A comparative study on transformer vs RNN in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, et al., · 2019
Cited alongside, same era.
“Transformer-XL: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov, · 2019
Cited alongside, same era.
“Transformer-transducer: End-to-end speech recognition with self-attention,”
Ching-Feng Yeh, Jay Mahadeokar, et al., · 2019
Cited alongside, same era.
“Improving layer trajectory lstm with future context frames,”
Jinyu Li, Liang Lu, Changliang Liu, and Yifan Gong, · 2019
Cited alongside, same era.
“Developing RNN-T models surpassing high-performance hybrid models with customization capability,”
Jinyu Li, , Rui Zhao, Zhong Meng, et al., · 2020
Cited alongside, same era.
Closest in time.
“Streaming automatic speech recognition with the transformer model,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2020
Closest in time.
“Universal ASR: Unify and improve streaming ASR with full-context modeling,”
Jiahui Yu, Wei Han, et al., · 2020
Closest in time.
“Transformer transducer: One model unifying streaming and non-streaming speech recognition,”
Anshuman Tripathi, Jaeyoung Kim, Qian Zhang, Han Lu, and Hasim Sak, · 2020
Closest in time.
“Synchronous transformers for end-to-end speech recognition,”
Zhengkun Tian, Jiangyan Yi, Ye Bai, Jianhua Tao, Shuai Zhang, and Zhengqi Wen, · 2020
Closest in time.
“Streaming transformer-based acoustic models using self-attention with augmented memory,”
Chunyang Wu, Yongqiang Wang, Yangyang Shi, Ching-Feng Yeh, and Frank Zhang, · 2020
Closest in time.
“Enhancing monotonic multihead attention for streaming asr,”
Hirofumi Inaguma, Masato Mimura, and Tatsuya Kawahara, · 2020
Closest in time.
“Transformer-based acoustic modeling for hybrid speech recognition,”
Yongqiang Wang, Abdelrahman Mohamed, et al., · 2020
Closest in time.