Fetching the paper…
Reading the bibliography…
The goal of this work is to develop a meeting transcription system that can recognize speech even when utterances of different speakers are overlapped.
J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small-room acoustics,”
1979
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, M. Mandel, and J. Le Roux, “Improved MVDR beamforming using single-channel mask prediction networks,” in
1985
Earlier work this paper cites.
J. G. Fiscus, N. Radde, J. S. Garofolo, J. A. A. Le, and C. Laprun, “The rich transcription 2005 spring meeting recogntion evaluation,” in
2005
Earlier work this paper cites.
O. Çetin and E. Shriberg, “Analysis of overlaps in meetings by dialog factors, hot spots, speakers, and collection site: Insights for automatic speech recognition,” in
2006
Earlier work this paper cites.
J. G. Fiscus, J. Ajot, N. Raddle, and C. Laprum, “Multiple dimension Levenshtein edit distance calculations for evaluating automatic speech recognition systems during simulaneous speech,” in
2006
Earlier work this paper cites.
M. Souden, J. Benesty, and S. Affes, “On optimal frequency-domain multichannel linear filtering for noise reduction,”
2007
Earlier work this paper cites.
X. Anguera, C. Wooters, and J. Hernando, “Acoustic beamforming for speaker diarization of meetings,”
2007
Earlier work this paper cites.
T. Hori, S. Araki, T. Yoshioka, M. Fujimoto, S. Watanabe, T. Oba, A. Ogawa, K. Otsuka, D. Mikami, K. Kinoshita, T. Nakatani, A. Nakamura, and J. Yamato, “Low-latency real-time meeting recognition and understanding using distant microphones and omni-directional camera,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi speech recognition toolkit,” in
2011
Earlier work this paper cites.
T. Hain, L. Burget, J. Dines, P. N. Garner, F. Grézl, A. El Hannani, M. Huijbregts, M. Karafiát, M. Lincoln, and V. Wan, “Transcribing meetings with the AMIDA systems,”
2012
Earlier work this paper cites.
T. Yoshioka and T. Nakatani, “Generalization of multi-channel linear prediction methods for blind MIMO impulse response shortening,”
2012
Cited alongside, same era.
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, “1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs,” in
2014
Cited alongside, same era.
2015
Cited alongside, same era.
A. Mohamed, F. Seide, D. Yu, J. Droppo, A. Stolcke, G. Zweig, and G. Penn, “Deep bi-directional recurrent networks over spectral windows,” in
2015
Cited alongside, same era.
2017
Later among the works it cites.
L. Drude and R. Haeb-Umbach, “Tight integration of spatial and spectral features for BSS with deep clustering embeddings,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
K. Žmolíková, M. Delcroix, K. Kinoshita, T. Higuchi, A. Ogawa, and T. Nakatani, “Learning speaker representation for neural network based multichannel speaker extraction,” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Yoshioka, N. Ito, M. Delcroix, A. Ogawa, K. Kinoshita, M. Fujimoto, C. Yu, W. J. Fabian, M. Espi, T. Higuchi, S. Araki, and T. Nakatani, “The NTT CHiME-3 system: advances in speech enhancement and recognition for mobile multi-microphone devices,” in
2015
Cited alongside, same era.
J. Heymann, L. Drude, A. Chinaev, and R. Haeb-Umbach, “BLSTM supported GEV beamformer front-end for the 3rd CHiME challenge,” in
2015
Cited alongside, same era.
2016
Cited alongside, same era.
J. R. Hershey, Z. Chen, J. L. Roux, and S. Watanabe, “Deep clustering: discriminative embeddings for segmentation and separation,” in
2016
Cited alongside, same era.
S. Renals and P. Swietojanski,
2017
Cited alongside, same era.
E. Edwards, W. Salloum, G. P. Finley, J. Fone, G. Cardiff, M. Miller, and D. Suendermann-Oeft, “Medical speech recognition: Reaching parity with humans,” in
2017
Cited alongside, same era.
2017
Later among the works it cites.
B. Li, T. N. Sainath, A. Narayanan, J. Caroselli, M. Bacchiani, A. Misra, I. Shafran, H. Sak, G. Punduk, K. Chin, K. C. Sim, R. J. Weiss, K. W. Wilson, E. Variani, C. Kim, O. Siohan, M. Weintrauba, E. McDermott, R. Rose, and M. Shannon, “Acoustic modeling for Google Home,” in
2017
Later among the works it cites.
J. Heymann, L. Drude, C. Boeddeker, P. Hanebrink, and R. Haeb-Umbach, “BeamNet: End-to-end training of a beamformer-supported multi-channel ASR system,” in
2017
Later among the works it cites.
T. Yoshioka, H. Erdogan, Z. Chen, and F. Alleva, “Multi-microphone neural speech separation for far-field multi-talker speech recognition,” in
2018
Closest in time.
C. Boeddeker, H. Erdogan, T. Yoshioka, and R. Haeb-Umbach, “Exploring practical aspects of neural mask-based beamforming for far-field speech recognition,” in
2018
Closest in time.
X. Zhang and D. Wang, “Binaural reverberant speech separation based on deep neural networks,” in
2022
Closest in time.