Fetching the paper…
Reading the bibliography…
Segmentation for continuous Automatic Speech Recognition (ASR) has traditionally used silence timeouts or voice activity detectors (VADs), which are both limited to acoustic features.
“Prosody-based automatic segmentation of speech into sentences and topics,”
Elizabeth Shriberg, Andreas Stolcke, Dilek Hakkani-Tür, and Gökhan Tür, · 2000
Earlier work this paper cites.
“Voice activity detection. fundamentals and speech recognition system robustness,”
Javier Ramirez, Juan Manuel Górriz, and José Carlos Segura, · 2007
Earlier work this paper cites.
“Formatting time-aligned asr transcripts for readability,”
Maria Shugrina, · 2010
Earlier work this paper cites.
“Nmt-based segmentation and punctuation insertion for real-time spoken language translation.,”
Eunah Cho, Jan Niehues, and Alex Waibel, · 2017
Earlier work this paper cites.
“Innovative method for unsupervised voice activity detection and classification of audio segments,”
Zulfiqar Ali and Muhammad Talha, · 2018
Earlier work this paper cites.
“Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,”
Roland Maas, Ariya Rastrow, Chengyuan Ma, Guitang Lan, Kyle Goehner, Gautam Tiwari, Shaun Joseph, and Björn Hoffmeister, · 2018
Earlier work this paper cites.
“Recognizing long-form speech using streaming end-to-end models,”
Arun Narayanan, Rohit Prabhavalkar, Chung-Cheng Chiu, David Rybach, Tara N Sainath, and Trevor Strohman, · 2019
Earlier work this paper cites.
“Joint endpointing and decoding with end-to-end models,”
Shuo-Yiin Chang, Rohit Prabhavalkar, Yanzhang He, Tara N Sainath, and Gabor Simko, · 2019
Earlier work this paper cites.
“Openwebtext corpus,” 2019
Aaron Gokaslan and Vanya Cohen, · 2019
Cited alongside, same era.
“End-to-end automatic speech recognition integrated with ctc-based voice activity detection,”
Takenori Yoshimura, Tomoki Hayashi, Kazuya Takeda, and Shinji Watanabe, · 2020
Cited alongside, same era.
“Segment boundary detection directed attention for online end-to-end speech recognition,”
Junfeng Hou, Wu Guo, Yan Song, and Li-Rong Dai, · 2020
Cited alongside, same era.
“Towards fast and accurate streaming end-to-end asr,”
Bo Li, Shuo-yiin Chang, Tara N Sainath, Ruoming Pang, Yanzhang He, Trevor Strohman, and Yonghui Wu, · 2020
Cited alongside, same era.
“End-to-end speech endpoint detection utilizing acoustic and language modeling knowledge for online low-latency speech recognition,”
Inyoung Hwang and Joon-Hyuk Chang, · 2020
Cited alongside, same era.
Meng Li, Shiyu Zhou, and Bo Xu, · 2021
Later among the works it cites.
“Vadoi: Voice-activity-detection overlapping inference for end-to-end long-form speech recognition,”
Jinhan Wang, Xiaosu Tong, Jinxi Guo, Di He, and Roland Maas, · 2022
Closest in time.
“Streaming punctuation for long-form dictation with transformers,”
Piyush Behre, Sharman Tan, Padma Varadharajan, and Shuangyu Chang, · 2022
Closest in time.
“Endpoint detection for streaming end-to-end multi-talker asr,”
Liang Lu, Jinyu Li, and Yifan Gong, · 2022
Closest in time.
“E2e segmenter: Joint segmenting and decoding for long-form asr,”
W Ronny Huang, Shuo-yiin Chang, David Rybach, Rohit Prabhavalkar, Tara N Sainath, Cyril Allauzen, Cal Peyser, and Zhiyun Lu, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiazheng Li, Linyi Yang, Barry Smyth, and Ruihai Dong, · 2020
Cited alongside, same era.
“Rnn-t models fail to generalize to out-of-domain audio: Causes and solutions,”
Chung-Cheng Chiu, Arun Narayanan, Wei Han, Rohit Prabhavalkar, Yu Zhang, Navdeep Jaitly, Ruoming Pang, Tara N Sainath, Patrick Nguyen, Liangliang Cao, et al., · 2021
Cited alongside, same era.
Zhiyun Lu, Yanwei Pan, Thibault Doutre, Liangliang Cao, Rohit Prabhavalkar, Chao Zhang, and Trevor Strohman, · 2021
Cited alongside, same era.
Closest in time.
Accessed: 2022-05-30
“Home page top stories,” https://www.npr.org/ , · 2022
Closest in time.
Accessed: 2022-05-30
“Debates and videos: Plenary: European parliament,” https://www.europarl.europa.eu/plenary/en/debates-video.html , · 2022
Closest in time.