Fetching the paper…
Reading the bibliography…
When we use End-to-end automatic speech recognition (E2E-ASR) system for real-world applications, a voice activity detection (VAD) system is usually needed to improve the performance and to reduce the computational cost by discarding non-speech parts in the audio.
B. Atal and L. Rabiner, “A pattern recognition approach to voiced-unvoiced-silence classification with applications to speech recognition,”
1976
Earlier work this paper cites.
J. Junqua, B. Reaves, and B. K. Mak, “A study of endpoint detection algorithms in adverse conditions: incidence on a dtw and hmm recognizer,” in
1991
Earlier work this paper cites.
R. Tucker, “Voice activity detection using a periodicity measure,” 1992
1992
Earlier work this paper cites.
P. Jusczyk, “How infants begin to extract words from speech,”
1999
Earlier work this paper cites.
K. H. Woo, T.-Y. Yang, K. J. Park, and C. Lee, “Robust voice activity detection algorithm for estimating noise spectrum,”
2000
Earlier work this paper cites.
A. Lee, K. Nakamura, R. Nisimura, H. Saruwatari, and K. Shikano, “Noise robust real world spoken dialogue system using gmm based rejection of unintended inputs,” in
2004
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
N. Mesgarani, M. Slaney, and S. Shamma, “Discrimination of speech from nonspeech based on multiscale spectro-temporal modulations,”
2006
Earlier work this paper cites.
Y. Liu, P. Fung, Y. Yang, C. Cieri, S. Huang, and D. Graff, “Hkust/mts: A very large scale mandarin telephone speech corpus,” in
2006
Earlier work this paper cites.
J. Ramirez, J. M. Górriz, and J. C. Segura, “Voice activity detection. fundamentals and speech recognition system robustness,” 2007
2007
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
X. Zhang and J. Wu, “Deep belief networks based voice activity detection,”
2013
Cited alongside, same era.
N. Ryant, M. Liberman, and J. Yuan, “Speech activity detection on youtube using deep neural networks,” in
2013
Cited alongside, same era.
T. Hughes and K. Mierle, “Recurrent neural networks for voice activity detection,”
2013
Cited alongside, same era.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in
2014
Cited alongside, same era.
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,”
L. Dong, F. Wang, and B. Xu, “Self-attention aligner: A latency-control end-to-end model for asr using self-attention network and chunk-hopping,”
2019
Later among the works it cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Later among the works it cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in
2019
Later among the works it cites.
V. Pratap, A. Hannun, Q. Xu, J. Cai, J. Kahn, G. Synnaeve, V. Liptchinsky, and R. Collobert, “Wav2letter++: A fast open-source speech recognition system,”
2019
Later among the works it cites.
T. Yoshimura, T. Hayashi, K. Takeda, and S. Watanabe, “End-to-end automatic speech recognition integrated with ctc-based voice activity detection,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
K. Rao, H. Sak, and R. Prabhavalkar, “Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
2017
Cited alongside, same era.
C. C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, and E. a. Gonina, “State-of-the-art speech recognition with sequence-to-sequence models,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in
2018
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Lavechin, M.-P. Gill, R. Bousbib, H. Bredin, and L. P. García-Perera, “End-to-end domain-adversarial voice activity detection,” in
2020
Later among the works it cites.
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M.-P. Gill, “pyannote.audio: neural building blocks for speaker diarization,” in
2020
Later among the works it cites.
F. Tao and C. Busso, “End-to-end audiovisual speech recognition system with multitask learning,”
2021
Closest in time.
S. Team, “Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier,” https://github.com/snakers4/silero-vad, 2021
2021
Closest in time.