Fetching the paper…
Reading the bibliography…
Stream fusion, also known as system combination, is a common technique in automatic speech recognition for traditional hybrid hidden Markov model approaches, yet mostly unexplored for modern deep neural network end-to-end model architectures.
D. B. Paul and J. M. Baker, “The Design for the Wall Street Journal-Based CSR Corpus,” in
1992
Earlier work this paper cites.
H. Bourlard and N. Morgan,
1994
Earlier work this paper cites.
J. G. Fiscus, “A Post-Processing System to Yield Reduced Word Error Rates: Recognizer Output Voting Error Reduction (ROVER),” in
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-Based Learning Applied to Document Recognition,”
1998
Earlier work this paper cites.
G. Potamianos, J. Luettin, and C. Neti, “Hierarchical Discriminant Features for Audio-Visual LVCSR,” in
2001
Earlier work this paper cites.
J. Luettin, G. Potamianos, and C. Neti, “Asynchronous Stream Modeling for Large Vocabulary Audio-Visual Speech Recognition,” in
2001
Earlier work this paper cites.
H. Misra, H. Bourlard, and V. Tyagi, “New Entropy Based Combination Rules in HMM/ANN Multi-Stream ASR,” in
2003
Earlier work this paper cites.
G. Potamianos
2003
Earlier work this paper cites.
R. Schlüter, A. Zolnay, and H. Ney, “ Feature Combination using Linear Discriminant Analysis and its Pitfalls,” in
2006
Earlier work this paper cites.
B. Hoffmeister, T. Klein, R. Schlüter, and H. Ney, “Frame Based System Combination and a Comparison With Weighted ROVER and CNC,” in
2006
Earlier work this paper cites.
H. Xu, D. Povey, L. Mangu, and J. Zhu, “Minimum Bayes Risk Decoding and System Combination Based on a Recursion for Edit Distance,”
2011
Earlier work this paper cites.
D. Povey
2011
Earlier work this paper cites.
A. Graves, A. Mohamed, and G. Hinton, “Speech Recognition with Deep Recurrent Neural Networks,” in
2013
Earlier work this paper cites.
E. Loweimi, S. M. Ahadi, and T. Drugman, “A New Phase-Based Feature Representation for Robust Speech Recognition,” in
2013
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards End-to-End Speech Recognition with Recurrent Neural Networks,” in
2014
Cited alongside, same era.
N. Srivastava
2014
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR Corpus Based on Public Domain Audio Books,” in
2015
Cited alongside, same era.
2015
Cited alongside, same era.
D. Bahdanau
2016
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition,” in
2016
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Lohrenz, W. Li, and T. Fingscheidt, “A New TIMIT Benchmark for Context-Independent Phone Recognition Using Turbo Fusion,” in
2018
Later among the works it cites.
T. Hori, J. Cho, and S. Watanabe, “End-to-End Speech Recognition With Word-Based RNN Language Models,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Receveur, R. Weiss, and T. Fingscheidt, “Turbo Automatic Speech Recognition,”
2016
Cited alongside, same era.
2017
Cited alongside, same era.
T. Lohrenz and T. Fingscheidt, “Turbo Fusion of Magnitude and Phase Information for DNN-Based Phoneme Recognition,” in
2017
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention Based End-to-End Speech Recognition Using Multi-task Learning,” in
2017
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition,” in
2018
Cited alongside, same era.
S. Karita
2019
Later among the works it cites.
M. Ravanelli, T. Parcollet, and Y. Bengio, “The Pytorch-Kaldi Speech Recognition Toolkit,” in
2019
Later among the works it cites.
D. S. Park
2019
Later among the works it cites.
A. Paszke
2019
Later among the works it cites.
E. Tsunoo, Y. Kashiwagi, T. Kumakura, and S. Watanabe, “Transformer ASR with Contextual Block Processing,” in
2019
Later among the works it cites.
S. Karita
2019
Later among the works it cites.
T. Lohrenz and T. Fingscheidt, “BLSTM-Driven Stream Fusion for Automatic Speech Recognition: Novel Methods and a Multi-Size Window Fusion Example,” in
2020
Later among the works it cites.
T. Moriya
2020
Later among the works it cites.
R. Müller, S. Kornblith, and G. Hinton, “When Does Label Smoothing Help?”
2020
Later among the works it cites.