Fetching the paper…
Reading the bibliography…
Large-scale, weakly-supervised speech recognition models, such as Whisper, have demonstrated impressive results on speech recognition across domains and languages.
F. Brugnara, D. Falavigna, and M. Omologo, “Automatic segmentation and labeling of speech based on hidden markov models,” Speech Communication , vol. 12, no. 4, pp. 357–370, 1993
1993
Earlier work this paper cites.
J. Godfrey and E. Holliman, “Switchboard-1 release 2 ldc97s62,” Linguistic Data Consortium , p. 34, 1993
1993
Earlier work this paper cites.
Y.-J. Kim and A. Conkie, “Automatic segmentation combining an hmm-based approach and spectral boundary correction,” in Seventh International conference on spoken language processing , 2002
2002
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal et al. , “The ami meeting corpus: A pre-announcement,” in Machine Learning for Multimodal Interaction: Second International Workshop, MLMI 2005, Edinburgh, UK, July 11-13, 2005, Revised Selected Papers 2 . Springer, 2006, pp. 28–39
2006
Earlier work this paper cites.
K. Gorman, J. Howell, and M. Wagner, “Prosodylab-aligner: A tool for forced alignment of laboratory speech,” Canadian Acoustics , vol. 39, no. 3, pp. 192–193, 2011
2011
Earlier work this paper cites.
J. Yuan, N. Ryant, M. Liberman, A. Stolcke, V. Mitra, and W. Wang, “Automatic phonetic segmentation using boundary models.” in Interspeech , 2013, pp. 2306–2310
2013
Earlier work this paper cites.
A. Stolcke, N. Ryant, V. Mitra, J. Yuan, W. Wang, and M. Liberman, “Highly accurate phonetic segmentation using boundary correction models and system fusion,” in ICASSP . IEEE, 2014, pp. 5552–5556
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi.” in Interspeech , vol. 2017, 2017, pp. 498–502
2017
Earlier work this paper cites.
G. Gelly and J.-L. Gauvain, “Optimization of rnn-based speech activity detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 3, pp. 646–656, 2018
2018
Cited alongside, same era.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Esteve, “Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in Speech and Computer: 20th International Conference, SPECOM 2018, Leipzig, Germany, September 18–22, 2018, Proceedings 20 . Springer, 2018, pp. 198–208
2018
Cited alongside, same era.
Y. Jia, M. Johnson, W. Macherey, R. J. Weiss, Y. Cao, C.-C. Chiu, N. Ari, S. Laurenzo, and Y. Wu, “Leveraging weakly supervised data to improve end-to-end speech-to-text translation,” in ICASSP . IEEE, 2019, pp. 7180–7184
2019
Cited alongside, same era.
C.-C. Chiu, W. Han, Y. Zhang, R. Pang, S. Kishchenko, P. Nguyen, A. Narayanan, H. Liao, S. Zhang, A. Kannan et al. , “A comparison of end-to-end models for long-form speech recognition,” in 2019 IEEE automatic speech recognition and understanding workshop (ASRU) . IEEE, 2019, pp. 889–896
2021
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Later among the works it cites.
J. Kang, J. Huh, H. S. Heo, and J. S. Chung, “Augmentation adversarial training for self-supervised speaker representation learning,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1253–1262, 2022
2022
Later among the works it cites.
Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass, “Ssast: Self-supervised audio spectrogram transformer,” in Proc. AAAI , vol. 36, no. 10, 2022, pp. 10 699–10 709
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” NeurIPS , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
S. Wisdom, E. Tzinis, H. Erdogan, R. Weiss, K. Wilson, and J. Hershey, “Unsupervised sound separation using mixture invariant training,” NeurIPS , vol. 33, pp. 3846–3857, 2020
2020
Cited alongside, same era.
L. Kürzinger, D. Winkelbauer, L. Li, T. Watzel, and G. Rigoll, “Ctc-segmentation of large corpora for german end-to-end speech recognition,” in Speech and Computer (SPECOM 2020) . Springer, 2020, pp. 267–278
2020
Cited alongside, same era.
2020
Cited alongside, same era.
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M.-P. Gill, “pyannote.audio: neural building blocks for speaker diarization,” in ICASSP , 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
Later among the works it cites.
2022
Later among the works it cites.
J. Li, Y. Meng, Z. Wu, H. Meng, Q. Tian, Y. Wang, and Y. Wang, “Neufa: Neural network based end-to-end forced alignment with bidirectional attention mechanism,” in ICASSP . IEEE, 2022, pp. 8007–8011
2022
Later among the works it cites.
Y.-Y. Yang, M. Hira, Z. Ni, A. Astafurov, C. Chen, C. Puhrsch, D. Pollack, D. Genzel, D. Greenberg, E. Z. Yang et al. , “Torchaudio: Building blocks for audio and speech processing,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6982–6986
2022
Later among the works it cites.
H. Chen, H. Zhang, L. Wang, K. A. Lee, M. Liu, and J. Dang, “Self-supervised audio-visual speaker representation with co-meta learning,” in ICASSP , 2023
2023
Closest in time.
J. Louradour, “whisper-timestamped,” https://github.com/linto-ai/whisper-timestamped/tree/f861b2b19d158f3cbf4ce524f22c78cb471d6131 , 2023
2023
Closest in time.
“Which automatic transcription service is the most accurate?” https://medium.com/descript/which-automatic-transcription-service-is-the-most-accurate-2018-2e859b23ed19, accessed: 2023-04-27
2023
Closest in time.