Fetching the paper…
Reading the bibliography…
Recently, self-attention-based transformers and conformers have been introduced as alternatives to RNNs for ASR acoustic modeling.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Capacity and trainability in recurrent neural networks,”
Jasmine Collins, Jascha Narain Sohl-Dickstein, and David Sussillo, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“AISHELL-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Earlier work this paper cites.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2018
Earlier work this paper cites.
“A comparative study on transformer vs rnn in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al., · 2019
Cited alongside, same era.
“Compressive transformers for long-range sequence modelling,”
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap, · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, Chung-Cheng Chiu, James Qin, Jiahui Yu, Niki Parmar, Ruoming Pang, Shibo Wang, Wei Han, Yonghui Wu, Yu Zhang, and Zhengdong Zhang, · 2020
Cited alongside, same era.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Cited alongside, same era.
“Transformers are RNNs: Fast autoregressive transformers with linear attention,”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret, · 2020
“Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,”
Yangyang Shi, Yongqiang Wang, Chunyang Wu, Ching feng Yeh, Julian Chan, Frank Zhang, Duc Le, and Michael L. Seltzer, · 2021
Later among the works it cites.
“Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,”
Guanbo Wang Guoguo Chen, Shuzhou Chai and et al., · 2021
Later among the works it cites.
“CUSIDE: Chunking, Simulating Future Context and Decoding for Streaming ASR,”
Keyu An, Huahuan Zheng, Zhijian Ou, Hongyu Xiang, Ke Ding, and Guanglu Wan, · 2022
Later among the works it cites.
“ConvRNN-T: Convolutional Augmented Recurrent Neural Network Transducers for Streaming Speech Recognition,”
Martin Radfar, Rohit Barnwal, Rupak Vignesh Swaminathan, Feng-Ju Chang, Grant P. Strimel, Nathan Susanj, and Athanasios Mouchtaris, · 2022
Later among the works it cites.
“Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition,”
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al., · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“CIF: Continuous integrate-and-fire for end-to-end speech recognition,”
Linhao Dong and Bo Xu, · 2020
Cited alongside, same era.
“Improving 1611.09913rnn transducer with normalized jointer network,”
Mingkun Huang, Jun Zhang, Meng Cai, Yang Zhang, Jiali Yao, Yongbin You, Yi He, and Zejun Ma, · 2020
Cited alongside, same era.
“Dual-mode ASR: Unify and improve streaming ASR with full-context modeling,”
Jiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu, Bo Li, Tara N Sainath, Yonghui Wu, and Ruoming Pang, · 2021
Cited alongside, same era.
“U2++: Unified two-pass bidirectional end-to-end model for speech recognition,”
Di Wu, Binbin Zhang, Chao Yang, Zhendong Peng, Wenjing Xia, Xiaoyu Chen, and Xin Lei, · 2021
Cited alongside, same era.
Later among the works it cites.
“Rwkv: Reinventing rnns for the transformer era,”
Bo Peng, Eric Alcaide, Quentin G. Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, G Kranthikiran, Xuming He, Haowen Hou, Przemyslaw Kazienko, Jan Kocoń, Jiaming Kong, Bartlomiej Koptyra, Hayden Lau, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Xiangru Tang, Bolun Wang, Johan Sokrates Wind, Stansilaw Wozniak, Ruichong Zhang, Zhenyuan Zhang, Qihang Zhao, Peng Zhou, Jian Zhu, and Rui Zhu, · 2023
Closest in time.
“Bat: Boundary aware transducer for memory-efficient and low-latency asr,”
Keyu An, Xian Shi, and Shiliang Zhang, · 2023
Closest in time.
“Funasr: A fundamental end-to-end speech recognition toolkit,”
Zhifu Gao, Zerui Li, Jiaming Wang, Haoneng Luo, Xian Shi, Mengzhe Chen, Yabin Li, Lingyun Zuo, Zhihao Du, Zhangyu Xiao, and Shiliang Zhang, · 2023
Closest in time.