Fetching the paper…
Reading the bibliography…
Automatic speech recognition (ASR) systems developed in recent years have shown promising results with self-attention models (e.g., Transformer and Conformer), which are replacing conventional recurrent neural networks.
“Spontaneous Speech Corpus of Japanese”
Kikuo Maekawa, Hanae Koiso, Sadaoki Furui and Hitoshi Isahara · 2000
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks”
Alex Graves and Navdeep Jaitly · 2014
Earlier work this paper cites.
“LibriSpeech: An ASR corpus based on public domain audio books”
Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition”
Tom Ko, Vijayaditya Peddinti, Daniel Povey and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition”
William Chan, Navdeep Jaitly, Quoc Le and Oriol Vinyals · 2016
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey Hinton · 2016
Earlier work this paper cites.
“Deep networks with stochastic depth”
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra and Kilian Weinberger · 2016
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John Hershey and Tomoki Hayashi · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Language modeling with gated convolutional networks”
Yann Dauphin, Angela Fan, Michael Auli and David Grangier · 2017
Earlier work this paper cites.
“Decoupled weight decay regularization”
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
“The LJ Speech Dataset”, https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
“Speech-Transformer: a no-recurrence sequence-to-sequence model for speech recognition”
Linhao Dong, Shuang Xu and Bo Xu · 2018
Cited alongside, same era.
“ESPnet: End-to-End Speech Processing Toolkit”
Shinji Watanabe et al · 2018
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions”
Jonathan Shen et al · 2018
Cited alongside, same era.
“Transformer-XL: Attentive Language Models beyond a Fixed-Length Context”
Zihang Dai et al · 2019
Cited alongside, same era.
“SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition”
Daniel Park et al · 2019
Cited alongside, same era.
“A comparative study on Transformer vs RNN in speech applications”
Shigeki Karita et al · 2019
Cited alongside, same era.
“Transformer transducer: A streamable speech recognition model with Transformer encoders and RNN-T loss”
Qian Zhang et al · 2020
Later among the works it cites.
“ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context”
Wei Han et al · 2020
Later among the works it cites.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis”
Jungil Kong, Jaehyeon Kim and Jaekyoung Bae · 2020
Later among the works it cites.
“Recent developments on espnet toolkit boosted by conformer”
Pengcheng Guo et al · 2021
Later among the works it cites.
“Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers”
Albert Gu et al · 2021
Later among the works it cites.
“A Comparative Study on Neural Architectures and Training Methods for Japanese Speech Recognition”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Neural speech synthesis with Transformer network”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao and Ming Liu · 2019
Cited alongside, same era.
“Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS”
Mutian He, Yan Deng and Lei He · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented Transformer for Speech Recognition”
Anmol Gulati et al · 2020
Cited alongside, same era.
“Transformers are RNNs: Fast autoregressive Transformers with linear attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2020
Cited alongside, same era.
“Big Bird: Transformers for longer sequences”
Manzil Zaheer et al · 2020
Cited alongside, same era.
“On position embeddings in BERT”
Benyou Wang et al · 2020
Cited alongside, same era.
Shigeki Karita, Yotaro Kubo, Michiel Bacchiani and Llion Jones · 2021
Later among the works it cites.
“ESPnet2-TTS: Extending the edge of TTS research”
Tomoki Hayashi et al · 2021
Later among the works it cites.
“Efficiently Modeling Long Sequences with Structured State Spaces”
Albert Gu, Karan Goel and Christopher Ré · 2022
Closest in time.
“It’s Raw! Audio Generation with State-Space Models”
Karan Goel, Albert Gu, Chris Donahue and Christopher Ré · 2022
Closest in time.
“SRU++: Pioneering Fast Recurrence with Attention for Speech Recognition”
Jing Pan, Tao Lei, Kwangyoun Kim, Kyu Han and Shinji Watanabe · 2022
Closest in time.
“Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding”
Yifan Peng, Siddharth Dalmia, Ian Lane and Shinji Watanabe · 2022
Closest in time.
“Self-Attention with Relative Position Representations”
Peter Shaw, Jakob Uszkoreit and Ashish Vaswani · 2074
Closest in time.