Fetching the paper…
Reading the bibliography…
Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization.
F. Tao and C. Busso, “Bimodal recurrent neural network for audiovisual voice activity detection,” in
1942
Earlier work this paper cites.
R. Maas, A. Rastrow, K. Goehner, G. Tiwari, S. Joseph, and B. Hoffmeister, “Domain-specific utterance end-point detection for speech recognition,” in
1947
Earlier work this paper cites.
J. L. Fleiss, “Measuring nominal scale agreement among many raters,”
1971
Earlier work this paper cites.
T. Ng, B. Zhang, L. Nguyen, S. Matsoukas, X. Zhou, N. Mesgarani, K. Vesely, and P. Matejka, “Developing a speech activity detection system for the darpa rats program,” in
1972
Earlier work this paper cites.
R. Chengalvarayan, “Robust energy normalization using speech/nonspeech discriminator for german connected digit recognition,” in
1999
Earlier work this paper cites.
W.-H. Shin, B.-S. Lee, Y.-K. Lee, and J.-S. Lee, “Speech/nonspeech classification using multiple features for robust endpoint detection,” in
2000
Earlier work this paper cites.
K.-H. Woo, T.-Y. Yang, K.-J. Park, and C. Lee, “Robust voice activity detection algorithm for estimating noise spectrum,”
2000
Earlier work this paper cites.
I. McCowan, J. Carletta, W. Kraaij, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, W. Post, D. Reidsma, and P. Wellner, “The AMI meeting corpus,” in
2005
Earlier work this paper cites.
X. Anguera, C. Wooters, and J. Hernando, “Robust speaker diarization for meetings: Icsi rt06s evaluation system,” in
2006
Earlier work this paper cites.
T. Petsatodis, A. Pnevmatikakis, and C. Boukis, “Voice activity detection using audio-visual information,” in
2009
Earlier work this paper cites.
S. Galliano, G. Gravier, and L. Chaubard, “The ESTER-2 evaluation campaign for the rich transcription of french radio broadcasts,” in
2009
Earlier work this paper cites.
J. W. Shin, J.-H. Chang, and N. S. Kim, “Voice activity detection based on statistical models and machine learning approaches,”
2010
Cited alongside, same era.
D. Dean, S. Sridharan, R. Vogt, and M. Mason, “The QUT-NOISE-TIMIT corpus for the evaluation of voice activity detection algorithms,” in
2010
Cited alongside, same era.
“The WebRTC project,” 2011. [Online]. Available: https://webrtc.org/
2011
Cited alongside, same era.
M. Zelenak, H. Schulz, and J. Hernando, “Speaker diarization of broadcast news in albayzin 2010 evaluation campaign,”
2012
Cited alongside, same era.
G. Gravier, G. Adda, N. Paulson, M. Carre, A. Giraudel, and O. Galibert, “The ETAPE corpus for the evaluation of speechbased tv content processing in the french language,” in
2012
Cited alongside, same era.
Y. Mroueh, E. Marcheret, and V. Goel, “Deep multimodal learning for audio-visual speech recognition,” in
2015
Later among the works it cites.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,” in
2015
Later among the works it cites.
M. Buchbinder, Y. Buchris, and I. Cohen, “Adaptive weighting parameter in audio-visual voice activity detection,” in
2016
Later among the works it cites.
J.-S. Chung and A. Zisserman, “Lip reading in the wild,” in
2016
Later among the works it cites.
I. Jang, C. Ahn, J. Seo, and Y. Jang, “Enhanced feature extraction for speech detection in media audio,” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Giraudel, M. Carre, V. Mapelli, J. Kahn, O. Galibert, and L. Quintard, “The REPERE corpus : a multimodal corpus for person recognition,” in
2012
Cited alongside, same era.
S. O. Sadjadi and J. H. L. Hansen, “Unsupervised speech activity detection using voicing measures and perceptual spectral flux,” in
2013
Cited alongside, same era.
N. Ryant, M. Liberman, and J. Yuan, “Speech activity detection on youtube using deep neural networks,” in
2013
Cited alongside, same era.
H. Ghaemmaghami, D. Dean, S. Kalantari, S. Sridharan, and C. Fookes, “Complete-linkage clustering for voice activity detection in audio and visual speech,” in
2015
Cited alongside, same era.
D. Dov, R. Talmon, and I. Cohen, “Audio-visual voice activity detection using diffusion maps,”
2015
Cited alongside, same era.
B. Lehner, G. Widmer, and R. Sonnleitner, “Improving voice activity detection in movies,” in
2015
Cited alongside, same era.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in
Cited in the paper.
S.-Y. Chang, B. Li, T. N. Sainath, G. Simko, and C. Parada, “Endpoint detection using grid long short-term memory networks for streaming speech recognition,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik, “AVA: A video dataset of spatio-temporally localized atomic visual actions,” in
2018
Closest in time.
K. Hoover, S. Chaudhuri, C. Pantofaru, M. Slaney, and I. Sturdy, “Putting a face to the voice: Fusing audio and visual signals across a video to determine speakers,” in
2018
Closest in time.