Fetching the paper…
Reading the bibliography…
We demonstrate that carefully adjusting the tokenizer of the Whisper speech recognition model significantly improves the precision of word-level timestamps when applying dynamic time warping to the decoder's cross-attention scores.
A. Romana, J. Bandon, M. Perez, S. Gutierrez, R. Richter, A. Roberts, and E. M. Provost, “Automatically detecting errors and disfluencies in read speech to predict cognitive impairment in people with parkinson’s disease.” in INTERSPEECH , 2021, pp. 1907–1911
1911
Earlier work this paper cites.
P. Gage, “A new algorithm for data compression,” C Users Journal , vol. 12, no. 2, pp. 23–38, 1994
1994
Earlier work this paper cites.
E. G. Bard, R. J. Lickley, and M. P. Aylett, “Is disfluency just difficulty?” in Proc. ITRW on Disfluency in Spontaneous Speech (DiSS 2001) , 2001, pp. 97–100
2001
Earlier work this paper cites.
H. H. Clark and J. E. F. Tree, “Using uh and um in spontaneous speaking,” Cognition , vol. 84, no. 1, pp. 73–111, 2002
2002
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal et al. , “The ami meeting corpus: A pre-announcement,” in International workshop on machine learning for multimodal interaction . Springer, 2005, pp. 28–39
2005
Earlier work this paper cites.
2006
Earlier work this paper cites.
A. Lindström, J. Villing, S. Larsson, A. Seward, N. Åberg, and C. Holtelius, “The effect of cognitive load on disfluencies during in-vehicle spoken dialogue,” INTERSPEECH, 2008 , pp. 1196–1199, 09 2008
2008
Earlier work this paper cites.
M. Corley and O. Stewart, “Hesitation disfluencies in spontaneous speech: The meaning of,” Language and Linguistics Compass , vol. 2, pp. 589–602, 07 2008
2008
Earlier work this paper cites.
T. Giorgino, “Computing and visualizing dynamic time warping alignments in r: The dtw package,” Journal of Statistical Software , vol. 31, no. 7, 2009
2009
Earlier work this paper cites.
B. MacWhinney, D. Fromm, M. Forbes, and A. Holland, “Aphasiabank: Methods for studying discourse,” Aphasiology , vol. 25, no. 11, pp. 1286–1307, 2011
2011
Earlier work this paper cites.
K. Womack, W. McCoy, C. Ovesdotter Alm, C. Calvelli, J. B. Pelz, P. Shi, and A. Haake, “Disfluencies as extra-propositional indicators of cognitive processing,” in Proceedings of the Workshop on Extra-Propositional Aspects of Meaning in Computational Linguistics , R. Morante and C. Sporleder, Eds. Jeju, Republic of Korea: Association for Computational Linguistics, Jul. 2012, pp. 1–9. [Online]. Available: https://aclanthology.org/W12-3801
2012
Earlier work this paper cites.
R. Ochshorn and M. M. Hawkins, “Gentle: A robust yet lenient forced aligner built on kaldi,” https://lowerquality.com/gentle/ , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on . IEEE, 2015, pp. 5206–5210
2015
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Cited alongside, same era.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 776–780
2017
Cited alongside, same era.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Estève, “Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in Speech and Computer . Springer International Publishing, 2018, pp. 198–208
2022
Later among the works it cites.
2023
Later among the works it cites.
C. Lea, Z. Huang, J. Narain, L. Tooley, D. Yee, D. T. Tran, P. Georgiou, J. P. Bigham, and L. Findlater, “From user perceptions to technical improvement: Enabling people who stutter to better use speech recognition,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , 2023, pp. 1–16
2023
Later among the works it cites.
L. Wagner, M. Zusag, and T. Bloder, “Careful whisper – leveraging advances in automatic speech recognition for robust and interpretable aphasia subtype classification,” in INTERSPEECH , 2023
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
E. Fonseca, M. Plakal, D. P. Ellis, F. Font, X. Favory, and X. Serra, “Learning sound event classifiers from web audio with noisy labels,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 21–25
2019
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020) , 2020, pp. 4211–4215
2020
Cited alongside, same era.
2021
Cited alongside, same era.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Cited alongside, same era.
Y. Zhang, D. S. Park, W. Han, J. Qin, A. Gulati, J. Shor, A. Jansen, Y. Xu, Y. Huang, S. Wang et al. , “Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1519–1532, 2022
2022
Cited alongside, same era.
J. Kang, J. Huh, H. S. Heo, and J. S. Chung, “Augmentation adversarial training for self-supervised speaker representation learning,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1253–1262, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Later among the works it cites.
M. Bain, J. Huh, T. Han, and A. Zisserman, “Whisperx: Time-accurate speech transcription of long-form audio,” INTERSPEECH , 2023
2023
Later among the works it cites.
J. W. Kim, “openai-dtw,” https://github.com/openai/whisper/blob/main/notebooks/Multilingual_ASR.ipynb , 2023
2023
Later among the works it cites.
J. Louradour, “whisper-timestamped,” https://github.com/linto-ai/whisper-timestamped , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Wollin-Giering, M. Hoffmann, J. Höfting, C. Ventzke et al. , “Automatic transcription of english and german qualitative interviews,” in Forum Qualitative Sozialforschung/Forum: Qualitative Social Research , vol. 25, no. 1, 2024
2024
Closest in time.
2024
Closest in time.
OpenAI, “Whisper large v2,” https://huggingface.co/openai/whisper-large-v2 , 2024, version of 20.02.2024
2024
Closest in time.