Fetching the paper…
Reading the bibliography…
In this paper, we present a new open source toolkit for automatic speech recognition (ASR), named CAT (CRF-based ASR Toolkit).
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukáš Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlíček, Yanmin Qian, Petr Schwarz, Jan Silovský, Georg Stemmer, and Karel Vesel, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Speaker adaptation of neural network acoustic models using i-vectors,”
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, · 2013
Earlier work this paper cites.
“RASR/NN: The RWTH neural network toolkit for speech recognition,”
S. Wiesler, A. Richard, P. Golik, R. Schlüter, and H. Ney, · 2014
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep rnn models and wfst-based decoding,”
Y. Miao, M. Gowayyed, and F. Metze, · 2015
Earlier work this paper cites.
“Deep speech 2 : End-to-end speech recognition in english and mandarin,”
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, and et al., · 2015
Earlier work this paper cites.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Purely sequence-trained neural networks for ASR based on lattice-free MMI,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Earlier work this paper cites.
“Training deep bidirectional lstm acoustic model for lvcsr by a context-sensitive-chunk bptt approach,”
Chen Kai and Huo Qiang, · 2016
Cited alongside, same era.
“Towards online-recognition with deep bidirectional LSTM acoustic models,”
Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2016
Cited alongside, same era.
“Automatic differentiation in PyTorch,”
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, · 2017
Cited alongside, same era.
“Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM,”
Takaaki Hori, Shinji Watanabe, Yu Zhang, and William Chan, · 2017
Cited alongside, same era.
“wav2letter++: The fastest open-source speech recognition system,”
Qiantong Xu Jeff Cai Jacob Kahn Gabriel Synnaeve Vitaliy Liptchinsky Ronan Collobert Vineel Pratap, Awni Hannun, · 2018
Cited alongside, same era.
“Low latency acoustic modeling using temporal convolution and LSTMs,”
V. Peddinti, Y. Wang, D. Povey, and S. Khudanpur, · 2018
Later among the works it cites.
“Pykaldi: A Python wrapper for Kaldi,”
Dogan Can, Victor R. Martinez, Pavlos Papadopoulos, and Shrikanth S. Narayanan, · 2018
Later among the works it cites.
“A pruned rnnlm lattice-rescoring algorithm for automatic speech recognition,”
H. Xu, T. Chen, D. Gao, Y. Wang, K. Li, N. Goel, Y. Carmiel, D. Povey, and S. Khudanpur, · 2018
Later among the works it cites.
“RETURNN as a generic flexible neural toolkit with application to translation and speech recognition,”
Albert Zeyer, Tamer Alkhouli, and Hermann Ney, · 2018
Later among the works it cites.
“CRF-based single-stage acoustic modeling with CTC topology,”
Hongyu Xiang and Zhijian Ou, · 2019
Closest in time.
“Rwth asr systems for librispeech: Hybrid vs attention,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai, · 2018
Cited alongside, same era.
“Flat-start single-stage discriminatively trained HMM-based models for ASR,”
Hossein Hadian, Hossein Sameti, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Cited alongside, same era.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2018
Cited alongside, same era.
“Fully convolutional speech recognition,”
Neil Zeghidour, Qiantong Xu, Vitaliy Liptchinsky, Nicolas Usunier, Gabriel Synnaeve, and Ronan Collobert, · 2018
Cited alongside, same era.
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Closest in time.
“Online hybrid ctc/attention architecture for end-to-end speech recognition,”
Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Ta Li, and Yonghong Yan, · 2019
Closest in time.
“The Pytorch-kaldi speech recognition toolkit,”
M. Ravanelli, T. Parcollet, and Y. Bengio, · 2019
Closest in time.
“End-to-end speech recognition with adaptive computation steps,”
M. Li, M. Liu, and H. Masanori, · 2019
Closest in time.
“Component fusion: Learning replaceable language model component for end-to-end speech recognition system,”
Changhao Shan, Chao Weng, Guangsen Wang, Dan Su, Min Luo, Dong Yu, and Lei Xie, · 2019
Closest in time.