Fetching the paper…
Reading the bibliography…
End-to-end automatic speech recognition directly maps input speech to characters.
“Corpus of spontaneous japanese: Its design and evaluation,”
Kikuo Maekawa, · 2003
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“End-to-end speech recognition for languages with ideographic characters,”
Hitoshi Ito, Aiko Hagiwara, Manon Ichiki, Takeshi Mishima, Shoei Sato, and Akio Kobayashi, · 2017
Earlier work this paper cites.
“Multitask Learning with Low-Level Auxiliary Tasks for Encoder-Decoder Based Speech Recognition,”
Shubham Toshniwal, Hao Tang, Liang Lu, and Karen Livescu, · 2017
Earlier work this paper cites.
“Multi-accent speech recognition with hierarchical grapheme based models,”
Kanishka Rao and Haşim Sak, · 2017
Earlier work this paper cites.
“AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Hierarchical Multitask Learning With CTC,”
Ramon Sanabria and Florian Metze, · 2018
Cited alongside, same era.
“ESPnet: End-to-End Speech Processing Toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai, · 2018
Cited alongside, same era.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“Non-Autoregressive Transformer for Speech Recognition,”
Nanxin Chen, Shinji Watanabe, Jesús Villalba, Piotr Żelasko, and Najim Dehak, · 2021
Later among the works it cites.
“Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment,”
Ethan A. Chi, Julian Salazar, and Katrin Kirchhoff, · 2021
Later among the works it cites.
“Intermediate Loss Regularization for CTC-Based Speech Recognition,”
Jaesong Lee and Shinji Watanabe, · 2021
Later among the works it cites.
“Relaxing the Conditional Independence Assumption of CTC-Based ASR by Conditioning on Intermediate Predictions,”
Jumon Nozaki and Tatsuya Komatsu, · 2021
Later among the works it cites.
“A comparative study on non-autoregressive modelings for speech-to-text generation,”
Yosuke Higuchi, Nanxin Chen, Yuya Fujita, Hirofumi Inaguma, Tatsuya Komatsu, Jaesong Lee, Jumon Nozaki, Tianzi Wang, and Shinji Watanabe, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Cited alongside, same era.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented Transformer for Speech Recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Cited alongside, same era.
“Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict,”
Yosuke Higuchi, Shinji Watanabe, Nanxin Chen, Tetsuji Ogawa, and Tetsunori Kobayashi, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Joint phoneme-grapheme model for end-to-end speech recognition,”
Yotaro Kubo and Michiel Bacchiani, · 2020
Cited alongside, same era.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Later among the works it cites.
“End-to-end ASR to jointly predict transcriptions and linguistic annotations,”
Motoi Omachi, Yuya Fujita, Shinji Watanabe, and Matthew Wiesner, · 2021
Later among the works it cites.
“Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword Units,”
Yosuke Higuchi, Keita Karube, Tetsuji Ogawa, and Tetsunori Kobayashi, · 2021
Later among the works it cites.
“Recent developments on espnet toolkit boosted by conformer,”
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, and Yuekai Zhang, · 2021
Later among the works it cites.
“An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition,”
Xuankai Chang, Takashi Maekaku, Pengcheng Guo, Jing Shi, Yen-Ju Lu, Aswin Shanmugam Subramanian, Tianzi Wang, Shu-wen Yang, Yu Tsao, Hung-yi Lee, and Shinji Watanabe, · 2021
Later among the works it cites.
“Improving ctc-based asr models with gated interlayer collaboration,”
Yuting Yang, Yuke Li, and Binbin Du, · 2022
Closest in time.