Understand
Transcription of broadcast news is an interesting and challenging application for large-vocabulary continuous speech recognition (LVCSR).
- We present in detail the structure of a manually segmented and annotated corpus including over 160 hours of German broadcast news, and propose it as an evaluation framework of LVCSR systems.
- We show our own experimental results on the corpus, achieved with a state-of-the-art LVCSR decoder, measuring the effect of different feature sets and decoding parameters, and thereby demonstrate that real-time decoding of our test set is feasible on a desktop PC at 9.2% word error rate.
Built on
D. Willett, C. Neukirchen, and G. Rigoll, “Ducoder - the Duisburg University LVCSR stackdecoder,” in
2000
Earlier work this paper cites.
G. Rigoll, “The ALERT-System: Advanced broadcast speech recognition technology for selective dissemination of multimedia information,” in
2001
Earlier work this paper cites.
P. Beyerlein, X. Aubert, R. Haeb-Umbach, M. Harris, D. Klakow, A. Wendemuth, S. Molau, H. Ney, M. Pitz, and A. Sixtus, “Large vocabulary continuous speech recognition of broadcast news – the Philips/RWTH approach,”
2002
Earlier work this paper cites.
J.-L. Gauvain, L. Lamel, and G. Adda, “The LIMSI broadcast news transcription system,”
2002
Earlier work this paper cites.
K. McTait and M. Adda-Decker, “The 300k LIMSI German broadcast news transcription system,” in
2003
Earlier work this paper cites.
Similar
H. Yu, Y.-C. Tam, T. Schaaf, S. Stüker, Q. Jin, M. Noamany, and T. Schultz, “The ISL RT04 Mandarin broadcast news evaluation system,” in
2004
Cited alongside, same era.
J.-L. Gauvain, G. Adda, M. Adda-Decker, A. Allauzen, V. Gendner, L. Lamel, and H. Schwenk, “Where are we in transcribing French broadcast news?” in
2005
Cited alongside, same era.
M. J. F. Gales, D. Kim, P. C. Woodland, H. Chan, D. Mrva, R. Sinha, and S. Tranter, “Progress in the CU-HTK broadcast news transcription system,”
2006
Cited alongside, same era.
R. Sinha, M. J. F. Gales, D. Kim, X. Liu, K. Sim, and P. C. Woodland, “The CU-HTK broadcast news transcription system,” in
2006
Cited alongside, same era.
Then
S. J. Young, G. Evermann, M. J. F. Gales, D. Kershaw, G. Moore, J. J. Odell, D. G. Ollason, D. Povey, V. Valtchev, and P. C. Woodland,
2006
Later among the works it cites.
L. Lamel, A. Messaoudi, and J.-L. Gauvain, “Improved acoustic modeling for transcribing Arabic broadcast data,” in
2007
Later among the works it cites.
F. Eyben, M. Wöllmer, and B. Schuller, “openEAR - introducing the Munich open-source Emotion and Affect Recognition toolkit,” in
2009
Later among the works it cites.
2010
Later among the works it cites.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…