Fetching the paper…
Reading the bibliography…
Filler words such as `uh' or `um' are sounds or words people use to signal they are pausing to think.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in
1992
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: A resource for the next generations of speech-to-text.” in
2004
Earlier work this paper cites.
P. Howell, S. Davis, and J. Bartrip, “The university college london archive of stuttered speech (uclass),”
2009
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Earlier work this paper cites.
K. Womack, W. McCoy, C. O. Alm, C. Calvelli, J. B. Pelz, P. Shi, and A. Haake, “Disfluencies as extra-propositional indicators of cognitive processing,” in
2012
Earlier work this paper cites.
R. Gupta, K. Audhkhasi, S. Lee, and S. S. Narayanan, “Paralinguistic event detection from speech using probabilistic time-series smoothing and masking.” in
2013
Earlier work this paper cites.
H. Salamin, A. Polychroniou, and A. Vinciarelli, “Automatic detection of laughter and fillers in spontaneous mobile phone conversations,” in
2013
Earlier work this paper cites.
H. Hassan, L. Schwartz, D. Hakkani-Tür, and G. Tur, “Segmentation and disfluency removal for conversational speech translation,” in
2014
Earlier work this paper cites.
J. Ferguson, G. Durrett, and D. Klein, “Disfluency detection with a semi-markov model and prosodic features,” in
2015
Earlier work this paper cites.
L. Kaushik, A. Sangwan, and J. H. Hansen, “Laughter and filler detection in naturalistic audio.” International Speech and Communication Association, 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in
2015
Cited alongside, same era.
P. A. Heeman, R. Lunsford, A. McMillin, and J. S. Yaruss, “Using Clinician Annotations to Improve Automatic Speech Recognition of Stuttered Speech,” in
2016
Cited alongside, same era.
M. I. Tanveer, R. Zhao, K. Chen, Z. Tiet, and M. E. Hoque, “Automanner: An automated interface for making public speakers aware of their mannerisms,” in
2016
Cited alongside, same era.
C. Veaux, J. Yamagishi, K. MacDonald
2016
Cited alongside, same era.
A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for polyphonic sound event detection,”
2016
Cited alongside, same era.
V. Zayats and M. Ostendorf, “Giving attention to the unexpected: Using prosody innovations in disfluency detection,” in
2019
Later among the works it cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised Pre-Training for Speech Recognition,” in
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Wang, W. Che, Q. Liu, P. Qin, T. Liu, and W. Y. Wang, “Multi-task self-supervised learning for disfluency detection,” in
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Salamon, D. MacConnell, M. Cartwright, P. Li, and J. P. Bello, “Scaper: A library for soundscape synthesis and augmentation,” in
2017
Cited alongside, same era.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in
2017
Cited alongside, same era.
F. Wang, W. Chen, Z. Yang, Q. Dong, S. Xu, and B. Xu, “Semi-supervised disfluency detection,” in
2018
Cited alongside, same era.
P. Jamshid Lou, P. Anderson, and M. Johnson, “Disfluency detection using auto-correlational neural networks,” in
2018
Cited alongside, same era.
S. Das, N. Gandhi, T. Naik, and R. Shilkrot, “Increase apparent public speaking fluency by speech augmentation,” in
2019
Cited alongside, same era.
N. Bach and F. Huang, “Noisy bilstm-based models for disfluency detection.” in
2019
Cited alongside, same era.
R. Ochshorn and M. Max Hawkins, “Gentle: A robust yet lenient forced aligner built on kaldi.”
Cited in the paper.
Later among the works it cites.
Y. Chen, H. Dinkel, M. Wu, and K. Yu, “Voice Activity Detection in the Wild via Weakly Supervised Sound Event Detection,” in
2020
Later among the works it cites.
C. Lea, V. Mitra, A. Joshi, S. Kajarekar, and J. P. Bigham, “Sep-28k: A dataset for stuttering event detection from podcasts with people who stutter,” in
2021
Later among the works it cites.
S. A. Sheikh, M. Sahidullah, F. Hirsch, and S. Ouni, “Stutternet: Stuttering detection using time delay neural network,” in
2021
Later among the works it cites.
T. Kourkounakis, A. Hajavi, and A. Etemad, “Fluentnet: End-to-end detection of stuttered speech disfluencies with deep learning,”
2021
Later among the works it cites.
G. Chen, S. Chai, G.-B. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Z. You, and Z. Yan, “GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio,” in
2021
Later among the works it cites.
S. Hershey, D. P. Ellis, E. Fonseca, A. Jansen, C. Liu, R. C. Moore, and M. Plakal, “The benefit of temporally-strong labels in audio event classification,” in
2021
Later among the works it cites.