Fetching the paper…
Reading the bibliography…
We present a complete training pipeline to build a state-of-the-art hybrid HMM-based ASR system on the 2nd release of the TED-LIUM corpus.
Connectionist Speech Recognition: A Hybrid Approach
Herve A. Bourlard and Nelson Morgan, · 1993
Earlier work this paper cites.
“Improved backing-off for m-gram language modeling,”
Reinhard Kneser and Hermann Ney, · 1995
Earlier work this paper cites.
“Long Short-Term Memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“An Empirical Study of Smoothing Techniques for Language Modeling,”
Stanley F Chen and Joshua Goodman, · 1999
Earlier work this paper cites.
“Hypothesis Spaces for Minimum Bayes Risk Training in Large Vocabulary Speech Recognition,”
Matthew Gibson and Thomas Hain, · 2006
Earlier work this paper cites.
“Greedy Layer-Wise Training of Deep Networks,”
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle, · 2007
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, · 2011
Earlier work this paper cites.
“On the Estimation of Discount Parameters for Language Model Smoothing,”
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney, · 2011
Earlier work this paper cites.
“LSTM Neural Networks for Language Modeling,”
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney, · 2012
Earlier work this paper cites.
“Enhancing the TED-LIUM Corpus with Selected Data for Language Modeling and More TED Talks,”
Anthony Rousseau, Paul Deléglise, and Yannick Estève, · 2014
Earlier work this paper cites.
“RASR/NN: The RWTH Neural Network Toolkit for Speech Recognition,”
Simon Wiesler, Alexander Richard, Pavel Golik, Ralf Schlüter, and Hermann Ney, · 2014
Cited alongside, same era.
“Dropout: A Simple Way to Prevent Neural Networks from Overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Cited alongside, same era.
“Incorporating Nesterov Momentum into Adam,”
Timothy Dozat, · 2016
Cited alongside, same era.
“Attention Is All You Need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“RETURNN: the RWTH extensible training framework for universal recurrent neural networks,”
Patrick Doetsch, Albert Zeyer, Paul Voigtlaender, Ilia Kulikov, Ralf Schlüter, and Hermann Ney, · 2017
Cited alongside, same era.
“A Comprehensive Study of Deep Bidirectional LSTM RNNs for Acoustic Modeling in Speech Recognition,”
“RWTH ASR Systems for LibriSpeech: Hybrid vs Attention,”
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“Cumulative Adaptation for BLSTM Acoustic Models,”
Markus Kitza, Pavel Golik, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“Language Modeling with Deep Transformers,”
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“LSTM Language Models for LVCSR in First-Pass Decoding and Lattice-Rescoring,”
Eugen Beck, Wei Zhou, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“SpecAugment: A Simple Augmentation Method for Automatic Speech Recognition,”
Barret Zoph, Chung-Cheng Chiu, Daniel S. Park, Ekin Dogus Cubuk, Quoc V. Le, William Chan, and Yu Zhang, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Albert Zeyer, Patrick Doetsch, Paul Voigtlaender, Ralf Schlüter, and Hermann Ney, · 2017
Cited alongside, same era.
“Focal Loss for Dense Object Detection,”
Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollár, · 2017
Cited alongside, same era.
“Sisyphus, a Workflow Manager Designed for Machine Translation and Automatic Speech Recognition,”
Jan-Thorsten Peter, Eugen Beck, and Hermann Ney, · 2018
Cited alongside, same era.
“The CAPIO 2017 Conversational Speech Recognition System,”
Kyu J. Han, Akshay Chandrashekaran, Jungsuk Kim, and Ian R. Lane, · 2018
Cited alongside, same era.
“SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“A comparison of Transformer and LSTM encoder decoder models for ASR,”
Albert Zeyer, Parnia Bahar, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“A Comparative Study on Transformer vs RNN in Speech Applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, Shinji Watanabe, Takenori Yoshimura, and Wangyou Zhang, · 2019
Later among the works it cites.
“On Using SpecAugment for End-to-End Speech Translation,”
Parnia Bahar, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“Multi-Stride Self-Attention for Speech Recognition,”
Kyu J. Han, Jing Huang, Yun Tang, Xiaodong He, and Bowen Zhou, · 2019
Later among the works it cites.
“How Much Self-attention Do We Need? Trading Attention for Feed-forward Layers,”
Kazuki Irie, Alexander Gerstenberger, Ralf Schlüter, and Hermann Ney, · 2020
Closest in time.