Fetching the paper…
Reading the bibliography…
We propose a method of segmenting long-form speech by separating semantically complete sentences within the utterance.
J. Li, D. Yu, J.-T. Huang, and Y. Gong, “Improving wideband speech recognition using mixed-bandwidth training data in cd-dnn-hmm,” in 2012 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2012, pp. 131–136
2012
Earlier work this paper cites.
R. Zazo Candil, T. N. Sainath, G. Simko, and C. Parada, “Feature learning with raw-waveform cldnns for voice activity detection,” 2016
2016
Earlier work this paper cites.
C. Kim, A. Misra, K. Chin et al. , “Generation of Large-Scale Simulated Utterances in Virtual Rooms to Train Deep-Neural Networks for Far-Field Speech Recognition in Google Home,” in Proc. Interspeech , 2017
2017
Earlier work this paper cites.
C. Attig, N. Rauh, T. Franke, and J. F. Krems, “System latency guidelines then and now–is zero latency really considered necessary?” in Engineering Psychology and Cognitive Ergonomics: Cognition and Design: 14th International Conference, EPCE 2017, Held as Part of HCI International 2017, Vancouver, BC, Canada, July 9-14, 2017, Proceedings, Part II 14 . Springer, 2017, pp. 3–14
2017
Earlier work this paper cites.
D. S. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” in Proc. Interspeech , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Hlubík, M. Španěl, M. Boháč, and L. Weingartová, “Inserting punctuation to asr output in a real-time production environment,” in Text, Speech, and Dialogue: 23rd International Conference, TSD 2020, Brno, Czech Republic, September 8–11, 2020, Proceedings . Springer, 2020, pp. 418–425
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 1, jun 2020
2020
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
R. Prabhavalkar, Y. He, D. Rybach, S. Campbell, A. Narayanan, T. Strohman, and T. N. Sainath, “Less is more: Improved rnn-t decoding using limited label context and path merging,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5659–5663
2021
Cited alongside, same era.
J. Švec, J. Lehečka, L. Šmídl, and P. Ircing, “Transformer-based automatic punctuation prediction and word casing reconstruction of the asr output,” in Text, Speech, and Dialogue: 24th International Conference, TSD 2021, Olomouc, Czech Republic, September 6–9, 2021, Proceedings 24 . Springer, 2021, pp. 86–94
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Zhou, T. Tan, and Y. Qian, “Punctuation prediction for streaming on-device speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7277–7281
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
J. Yu, C.-C. Chiu, B. Li, S.-y. Chang, T. N. Sainath, Y. He, A. Narayanan, W. Han, A. Gulati, Y. Wu et al. , “Fastemit: Low-latency streaming asr with sequence-level emission regularization,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6004–6008
2021
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
E. Guzman, R. Botros, R. David, T. N. Sainath, W. Li, and Y. R. He, “Tied and reduced rnn-t decoder.”
Cited in the paper.
2022
Later among the works it cites.
2023
Closest in time.