Fetching the paper…
Reading the bibliography…
Is pushing numbers on a single benchmark valuable in automatic speech recognition? Research results in acoustic modeling are typically evaluated based on performance on a single dataset.
Communication in the presence of noise
C. E. Shannon · 1949
Earlier work this paper cites.
CSR-I (WSJ0) complete LDC93S6A
J. Garofolo, D. Graff, D. Paul, and D. Pallett · 1993
Earlier work this paper cites.
Switchboard-1 release 2 LDC97S62
J. Godfrey and E. Holliman · 1993
Earlier work this paper cites.
CSR-II (WSJ1) complete LDC94S13A
LDC and M. I. G. NIST · 1994
Earlier work this paper cites.
Large vocabulary continuous speech recognition using HTK
P. C. Woodland, J. J. Odell, V. Valtchev, and S. J. Young · 1994
Earlier work this paper cites.
2000 hub5 english evaluation speech LDC2002S09 and transcripts LDC2002T43
LDC et al · 2002
Earlier work this paper cites.
Fisher english training speech parts 1 and 2 transcripts LDC200{4,5}T19
C. Cieri, D. Graff, O. Kimball, D. Miller, and K. Walker · 2004
Earlier work this paper cites.
Fisher english training speech parts 1 and 2 LDC200{4,5}S13
C. Cieri, D. Miller, and K. Walker · 2004
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
2003 nist rich transcription evaluation data LDC2007S10
J. G. Fiscus et al · 2007
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
K. Heafield · 2011
Earlier work this paper cites.
The Kaldi speech recognition toolkit
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al · 2011
Earlier work this paper cites.
Revisiting recurrent neural networks for robust ASR
O. Vinyals, S. V. Ravuri, and D. Povey · 2012
Earlier work this paper cites.
An investigation of deep neural networks for noise robust speech recognition
M. L. Seltzer, D. Yu, and Y. Wang · 2013
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
D. Amodei et al · 2016
Cited alongside, same era.
Wav2letter: an end-to-end convnet-based speech recognition system
R. Collobert, C. Puhrsch, and G. Synnaeve · 2016
Cited alongside, same era.
The RWTH/UPB/FORTH system combination for the 4th CHiME challenge evaluation
T. Menne et al · 2016
Cited alongside, same era.
Domain adaptation of dnn acoustic models using knowledge distillation
T. Asami, R. Masumura, Y. Yamaguchi, H. Masataki, and Y. Aono · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Cited alongside, same era.
Investigation of transfer learning for ASR using LF-MMI trained neural networks
Adaptive input representations for neural language modeling
A. Baevski and M. Auli · 2019
Later among the works it cites.
SpecAugment: A simple data augmentation method for automatic speech recognition
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le · 2019
Later among the works it cites.
wav2letter++: The fastest open-source speech recognition system
V. Pratap, A. Hannun, Q. Xu, J. Cai, J. Kahn, G. Synnaeve, V. Liptchinsky, and R. Collobert · 2019
Later among the works it cites.
End-to-end ASR: from supervised to semi-supervised learning with modern architectures
G. Synnaeve, Q. Xu, J. Kahn, T. Likhomanenko, E. Grave, V. Pratap, A. Sriram, V. Liptchinsky, and R. Collobert · 2019
Later among the works it cites.
Building and evaluation of a real room impulse response dataset
I. Szöke, M. Skácel, L. Mošner, J. Paliesek, and J. Černockỳ · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Ghahremani, V. Manohar, H. Hadian, D. Povey, and S. Khudanpur · 2017
Cited alongside, same era.
The CAPIO 2017 conversational speech recognition system
K. J. Han, A. Chandrashekaran, J. Kim, and I. Lane · 2017
Cited alongside, same era.
A study on data augmentation of reverberant speech for robust speech recognition
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur · 2017
Cited alongside, same era.
Transfer learning for speech recognition on a budget
J. Kunze, L. Kirsch, I. Kurenkov, A. Krug, J. Johannsmeier, and S. Stober · 2017
Cited alongside, same era.
JHU Kaldi system for Arabic MGB-3 ASR challenge using diarization, audio-transcript alignment and transfer learning
V. Manohar, D. Povey, and S. Khudanpur · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
The fifth CHiME speech separation and recognition challenge: dataset, task and baselines
J. Barker, S. Watanabe, E. Vincent, and J. Trmal · 2018
Cited alongside, same era.
A. Andrusenko, A. Laptev, and I. Medennikov · 2020
Closest in time.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber · 2020
Closest in time.
Improving noise robustness of an end-to-end neural model for automatic speech recognition, 2020
J. Balam, J. Huang, V. Lavrukhin, S. Deng, S. Majumdar, and B. Ginsburg · 2020
Closest in time.
Reducing transformer depth on demand with structured dropout
A. Fan, E. Grave, and A. Joulin · 2020
Closest in time.
Conformer: Convolution-augmented transformer for speech recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al · 2020
Closest in time.
Whamr!: Noisy and reverberant single-channel speech separation
M. Maciejewski, G. Wichern, E. McQuinn, and J. Le Roux · 2020
Closest in time.
The stc system for the chime-6 challenge
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, et al · 2020
Closest in time.
Wer we are and wer we think we are
Szymański, Piotr, P. Żelasko, M. Morzy, A. Szymczak, M. Żyła-Hoppe, J. Banaszczak, L. Augustyniak, J. Mizgajski, and Y. Carmiel · 2020
Closest in time.
Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings
S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V. Manohar, D. Povey, D. Raj, et al · 2020
Closest in time.
Ccnet: Extracting high quality monolingual datasets from web crawl data
G. Wenzek, M.-A. Lachaux, A. Conneau, V. Chaudhary, F. Guzmán, A. Joulin, and É. Grave · 2020
Closest in time.
The RWTH ASR system for TED-LIUM release 2: Improving hybrid HMM with specaugment
W. Zhou, W. Michel, K. Irie, M. Kitza, R. Schlüter, and H. Ney · 2020
Closest in time.