Fetching the paper…
Reading the bibliography…
Automatic meeting analysis is an essential fundamental technology required to let, e.g.
CSR-I (WSJ0) Complete LDC93s6a
J. Garofolo, D. Graff, P. Doug, and D. Pallett, · 1993
Earlier work this paper cites.
“The AMI meeting corpus: A pre-announcement,”
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, I. McCowan, W. Post, D. Reidsma, , and P. Wellner, · 2006
Earlier work this paper cites.
“Performance measurement in blind audio source separation,”
E. Vincent, R. Gribonval, and C. Fevotte, · 2006
Earlier work this paper cites.
“Blind speech separation in a meeting situation with maximum SNR beamformers,”
S. Araki, H. Sawada, and S. Makino, · 2007
Earlier work this paper cites.
“Spring 2007 (rt-07) rich transcription meeting recognition evaluation plan,” 2007
NIST Speech Group, · 2007
Earlier work this paper cites.
“A DOA based speaker diarization system for real meetings,”
S. Araki, M. Fujimoto, K. Ishizuka, H. Sawada, and S. Makino, · 2008
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, , and P. Ouellet, · 2011
Earlier work this paper cites.
“Speaker diarization: A review of recent research,”
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, · 2012
Earlier work this paper cites.
“Low-latency real-time meeting recognition and understanding using distant microphones and omni-directional camera,”
T. Hori, S. Araki, T. Yoshioka, M. Fujimoto, S. Watanabe, T. Oba, A. Ogawa, K. Otsuka, D. Mikami, K. Kinoshita, T. Nakatani, A. Nakamura, and J. Yamato, · 2012
Earlier work this paper cites.
“Deep neural network-based speaker embeddings for end-to-end speaker verification,”
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, , and S. Khudanpur, · 2016
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
J. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, · 2016
Cited alongside, same era.
“Adaptive computation time for recurrent neural networks,” 2016,
A. Graves, · 2016
Cited alongside, same era.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
D. Yu, M. Kolbæk, Z. Tan, and J. Jensen, · 2017
Cited alongside, same era.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
M. Kolbæk, D. Yu, Z. Tan, and J. Jensen, · 2017
Cited alongside, same era.
“Deep Speaker: an end-to-end neural speaker embedding system,” 2017,
C. Li, X. Ma, B. Jiang, X. Li, X. Zhang, X. Liu, Y. Cao, A. Kannan, and Z. Zhu, · 2017
Cited alongside, same era.
“Dual frequency- and block-permutation alignment for deep learning based block-online blind source separation,”
L. Drude, T. Higuchi, K. Kinoshita, T. Nakatani, and R. Haeb-Umbach, · 2018
Later among the works it cites.
“Listening to each speaker one by one with recurrent selective hearing networks,”
K. Kinoshita, L. Drude, M. Delcroix, and T. Nakatani, · 2018
Later among the works it cites.
“Multi-microphone neural speech separation for far-field multi-talker speech recognition,”
T. Yoshioka, H. Erdogan, Z. Chen, and F. Alleva, · 2018
Later among the works it cites.
“End-to-end neural speaker diarization with permutation-free objectives,”
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, · 2019
Later among the works it cites.
“All-neural online source separation, counting, and diarization for meeting analysis,”
T. von Neumann and S. Araki T. Nakatani R. Haeb-Umbach K. Kinoshita, M. Delcroix, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Online meeting recognition in noisy environments with time-frequency mask based mvdr beamforming,”
S. Araki et al., · 2017
Cited alongside, same era.
First DIHARD Challenge Evaluation Plan
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, · 2018
Cited alongside, same era.
“BUT system for DIHARD speech diarization challenge 2018,”
M. Diez, F. Landini, L. Burget, J. Rohdin, A. Silnova, K. Zmolikova, O. Novotný, K. Veselý, O. Glembek, O. Plchot, L. Mošner, and P. Matějka, · 2018
Cited alongside, same era.
“Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,”
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, · 2018
Cited alongside, same era.
http://www.kecl.ntt.co.jp/icl/signal/kinoshita/publications/ ICASSP20/onlineRSAN/index.html
Cited in the paper.
“Recursive speech separation for unknown number of speakers,”
N. Takahashi, S. Parthasaarathy, N. Goswami, and Y. Mitsufuji, · 2019
Later among the works it cites.
“Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Later among the works it cites.
https://github.com/iiscleap/DIHARD-2019-baseline
2019
Later among the works it cites.
“Improving speaker discrimination of target speech extraction with time-domain speakerbeam,”
M. Delcroix, T. Ochiai, K. Zmolikova, K. Kinoshita, N. Tawara, T. Nakatani, and S. Araki, · 2020
Closest in time.