Fetching the paper…
Reading the bibliography…
In this study, we present recent developments on ESPnet: End-to-End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Performance measurement in blind audio source separation,”
Emmanuel Vincent, Rémi Gribonval, and Cédric Févotte, · 2006
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, et al., · 2011
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, et al., · 2017
Earlier work this paper cites.
“Language modeling with gated convolutional networks,”
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier, · 2017
Earlier work this paper cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Earlier work this paper cites.
“Speech-Transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Earlier work this paper cites.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, et al., · 2018
Cited alongside, same era.
“Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“End-to-end speech recognition with word-based RNN language models,”
Takaaki Hori, Jaejin Cho, and Shinji Watanabe, · 2018
Cited alongside, same era.
“Transformer-XL: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, et al., · 2019
Cited alongside, same era.
“A comparison of Transformer and LSTM encoder decoder models for ASR,”
Albert Zeyer, Parnia Bahar, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2019
Cited alongside, same era.
“A comparative study on Transformer vs RNN in speech applications,”
“Super-convergence: Very fast training of neural networks using large learning rates,”
Leslie N Smith and Nicholay Topin, · 2019
Later among the works it cites.
“Neural speech synthesis with Transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, et al., · 2020
Closest in time.
“ESPnet-ST: All-in-one speech translation toolkit,”
Hirofumi Inaguma, Shun Kiyono, Kevin Duh, Shigeki Karita, Nelson Enrique Yalta Soplin, et al., · 2020
Closest in time.
“ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,”
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, et al., · 2019
Cited alongside, same era.
“Learning deep Transformer models for machine translation,”
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, et al., · 2019
Cited alongside, same era.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, et al., · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, et al., · 2019
Cited alongside, same era.
“Improving Transformer-based end-to-end speech recognition with connectionist temporal classification and language model integration,”
Shigeki Karita, Nelson Enrique Yalta Soplin, Shinji Watanabe, Marc Delcroix, Atsunori Ogawa, et al., · 2019
Cited alongside, same era.
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler, · 2020
Closest in time.
“Understanding and improving Transformer from a multi-particle dynamic system point of view,”
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, et al., · 2020
Closest in time.
“FastSpeech 2: Fast and high-quality end-to-end text-to-speech,”
Yi Ren, Chenxu Hu, Tao Qin, Sheng Zhao, Zhou Zhao, et al., · 2020
Closest in time.
“Fastpitch: Parallel text-to-speech with pitch prediction,”
Adrian Lańcucki, · 2020
Closest in time.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim, · 2020
Closest in time.