Fetching the paper…
Reading the bibliography…
We describe a system that generates speaker-annotated transcripts of meetings by using a virtual microphone array, a set of spatially distributed asynchronous recording devices such as laptops and mobile phones.
H. Erdogan, J. R. Hershey, S. Watanabe, M. Mandel, and J. Le Roux, “Improved MVDR beamforming using single-channel mask prediction networks,” in
1985
Earlier work this paper cites.
J. G. Fiscus, “A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (ROVER),” in
1997
Earlier work this paper cites.
S. Tibrewala and H. Hermansky, “Sub-band based recognition of noisy speech,” in
1997
Earlier work this paper cites.
J. Gauvain, L. Lamel, and G. Adda, “Partitioning and transcription of broadcast news data,” in
1998
Earlier work this paper cites.
A. Stolcke, H. Bratt, J. Butzberger, H. Franco, V. R. Rao Gadde, M. Plauché, C. Richey, E. Shriberg, K. Sönmez, F. Weng, and J. Zheng, “The SRI March 2000 Hub-5 conversational speech transcription system,” in
2000
Earlier work this paper cites.
G. Evermann and P. Woodland, “Posterior probability decoding, confidence estimation, and system combination,” in
2000
Earlier work this paper cites.
A. Stolcke, “SRILM—an extensible language modeling toolkit,” in
2002
Earlier work this paper cites.
S. E. Tranter and D. A. Reynolds, “An overview of automatic speaker diarization systems,”
2006
Earlier work this paper cites.
J. G. Fiscus, J. Ajot, N. Raddle, and C. Laprum, “Multiple dimension Levenshtein edit distance calculations for evaluating automatic speech recognition systems during simulaneous speech,” in
2006
Earlier work this paper cites.
J. G. Fiscus, J. Ajot, and J. S. Garofolo, “The Rich Transcription 2007 meeting recognition evaluation,” in
2008
Earlier work this paper cites.
A. Stolcke, K. Boakye, Özgür Çetin, A. Janin, M. Magimai-Doss, C. Wooters, and J. Zheng, “The SRI-ICSI Spring 2007 meeting and lecture recognition system,” in
2008
Earlier work this paper cites.
R. Stiefelhagen, R. Bowers, and J. Fiscus, Eds.,
2008
Cited alongside, same era.
A. Stolcke, “Making the most from multiple microphones in meeting recordings,” in
2011
Cited alongside, same era.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in
2014
Cited alongside, same era.
T. Yoshioka, N. Ito, M. Delcroix, A. Ogawa, K. Kinoshita, M. Fujimoto, C. Yu, W. J. Fabian, M. Espi, T. Higuchi, S. Araki, and T. Nakatani, “The NTT CHiME-3 system: advances in speech enhancement and recognition for mobile multi-microphone devices,” in
2015
Cited alongside, same era.
J. Du, Y. Tu, L. Sun, F. Ma, H. Wang, J. Pan, C. Liu, J. Chen, and C. Lee, “The USTC-iFlytek system for CHiME-4 challenge,” in
2016
Cited alongside, same era.
S. Xue and Z. Yan, “Improving latency-controlled BLSTM acoustic models for online speech recognition,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Later among the works it cites.
J. Li, R. Zhao, Z. Chen, C. Liu, X. Xiao, G. Ye, and Y. Gong, “Developing far-field speaker system via teacher-student learning,” in
2018
Later among the works it cites.
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, “Diarization is hard: some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Heymann, L. Drude, and R. Haeb-Umbach, “Neural network based spectral mask estimation for acoustic beamforming,” in
2016
Cited alongside, same era.
T. Higuchi, T. Yoshioka, N. Ito, and T. Nakatani, “Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise,” in
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
2017
Cited alongside, same era.
B. Li, T. N. Sainath, A. Narayanan, J. Caroselli, M. Bacchiani, A. Misra, I. Shafran, H. Sak, G. Punduk, K. Chin, K. C. Sim, R. J. Weiss, K. W. Wilson, E. Variani, C. Kim, O. Siohan, M. Weintraub, E. McDermott, R. Rose, and M. Shannon, “Acoustic modeling for Google Home,” in
2017
Cited alongside, same era.
2018
Later among the works it cites.
S. Araki, N. Ono, K. Kinoshita, and M. Delcroix, “Meeting recognition with asynchronous distributed microphone array using block-wise refinement of mask-based MVDR beamformer,” in
2018
Later among the works it cites.
C. Boeddeker, H. Erdogan, T. Yoshioka, and R. Haeb-Umbach, “Exploring practical aspects of neural mask-based beamforming for far-field speech recognition,” in
2018
Later among the works it cites.
T. Yoshioka, H. Erdogan, Z. Chen, X. Xiao, and F. Alleva, “Recognizing overlapped speech in meetings: A multichannel separation approach using neural networks,” in
2018
Later among the works it cites.
T. Yoshioka, D. Dimitriadis, A. Stolcke, W. Hinthorn, Z. Chen, M. Zeng, and X. Huang, “Meeting transcription using asynchronous distant microphones,” in
2019
Closest in time.
A. Zhang, Q. Wan, Z. Zhu, J. Paisley, and C. Wang, “Fully supervised speaker diarization,” in
2019
Closest in time.