Fetching the paper…
Reading the bibliography…
Transformer has shown promising results in many sequence to sequence transformation tasks recently.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“End-to-end continuous speech recognition using attention-based recurrent NN: First results,”
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Improving neural machine translation models with monolingual data,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2015
Earlier work this paper cites.
“Deep speech 2: End-to-end speech recognition in English and Mandarin,”
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al., · 2016
Earlier work this paper cites.
“Online segment to segment neural transduction,”
Lei Yu, Jan Buys, and Phil Blunsom, · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Learning variable length units for SMT between related languages via byte pair encoding,”
Anoop Kunchukuttan and Pushpak Bhattacharyya, · 2016
Earlier work this paper cites.
“Purely sequence-trained neural networks for asr based on lattice-free mmi.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“Exploring neural transducers for end-to-end speech recognition,”
Eric Battenberg, Jitong Chen, Rewon Child, Adam Coates, Yashesh Gaur Yi Li, Hairong Liu, Sanjeev Satheesh, Anuroop Sriram, and Zhenyao Zhu, · 2017
Cited alongside, same era.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Cited alongside, same era.
“Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping.,”
Hasim Sak, Matt Shannon, Kanishka Rao, and Françoise Beaufays, · 2017
Cited alongside, same era.
“Attention is all you need,” 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Music transformer: Generating music with long-term structure,”
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck, · 2018
Later among the works it cites.
“Self-attention with relative position representations,”
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani, · 2018
Later among the works it cites.
“A comparative study on transformer vs rnn in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al., · 2019
Closest in time.
“The speechtransformer for large-scale Mandarin Chinese speech recognition,”
Jie Li, Xiaorui Wang, Yan Li, et al., · 2019
Closest in time.
“Parallel scheduled sampling,”
Daniel Duckworth, Arvind Neelakantan, Ben Goodrich, Lukasz Kaiser, and Samy Bengio, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Extending recurrent neural aligner for streaming end-to-end speech recognition in Mandarin,”
Linhao Dong, Shiyu Zhou, Wei Chen, and Bo Xu, · 2018
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Syllable-based sequence-to-sequence speech recognition with the transformer in Mandarin Chinese,”
Shiyu Zhou, Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“A comparison of modeling units in sequence-to-sequence speech recognition with the transformer on Mandarin Chinese,”
Shiyu Zhou, Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
Closest in time.
“Generating long sequences with sparse transformers,”
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever, · 2019
Closest in time.
“Transformer-XL: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov, · 2019
Closest in time.
“Recognizing long-form speech using streaming end-to-end models,”
Arun Narayanan, Rohit Prabhavalkar, Chung-Cheng Chiu, David Rybach, Tara N Sainath, and Trevor Strohman, · 2019
Closest in time.
“An online attention-based model for speech recognition,”
Ruchao Fan, Pan Zhou, Wei Chen, Jia Jia, and Gang Liu, · 2019
Closest in time.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Closest in time.