Fetching the paper…
Reading the bibliography…
Recently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis.
Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks
Morten Kolbaek, Dong Yu, Zheng-Hua Tan, and Jesper Jensen. 2017 · 1913
Earlier work this paper cites.
The AMI meeting corpus: A pre-announcement
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Maël Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner. 2005 · 2005
Earlier work this paper cites.
The rich transcription 2006 spring meeting recognition evaluation
Jonathan G. Fiscus, Jerome Ajot, Martial Michel, and John S. Garofolo. 2006 · 2006
Earlier work this paper cites.
A spectral clustering approach to speaker diarization
Huazhong Ning, Ming Liu, Hao Tang, and Thomas S. Huang. 2006 · 2006
Earlier work this paper cites.
Stream-based speaker segmentation using speaker factors and eigenvoices
Fabio Castaldo, Daniele Colibro, Emanuele Dalmasso, Pietro Laface, and Claudio Vair. 2008 · 2008
Earlier work this paper cites.
Front-end factor analysis for speaker verification
Najim Dehak, Patrick Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet. 2011 · 2011
Earlier work this paper cites.
MUSAN: A Music, Speech, and Noise Corpus
David Snyder, Guoguo Chen, and Daniel Povey. 2015 · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Developing on-line speaker diarization system
Dimitrios Dimitriadis and Petr Fousek. 2017 · 2017
Earlier work this paper cites.
Speaker diarization using deep neural network embeddings
Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, and Alan McCree. 2017 · 2017
Earlier work this paper cites.
A study on data augmentation of reverberant speech for robust speech recognition
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L. Seltzer, and Sanjeev Khudanpur. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines
Jon Barker, Shinji Watanabe, Emmanuel Vincent, and Jan Trmal. 2018 · 2018
Cited alongside, same era.
X-vectors: Robust DNN embeddings for speaker recognition
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. 2018 · 2018
Cited alongside, same era.
Generalized end-to-end loss for speaker verification
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez-Moreno. 2018 · 2018
Cited alongside, same era.
Deep-fsmn for large vocabulary continuous speech recognition
Shiliang Zhang, Ming Lei, Zhijie Yan, and Lirong Dai. 2018 · 2018
Cited alongside, same era.
Arcface: Additive angular margin loss for deep face recognition
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. 2019 · 2019
Cited alongside, same era.
Bayesian HMM based x-vector clustering for speaker diarization
Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds
Keisuke Kinoshita, Marc Delcroix, and Naohiro Tawara. 2021 · 2021
Later among the works it cites.
Discriminative neural clustering for speaker diarisation
Qiujia Li, Florian L. Kreyssig, Chao Zhang, and Philip C. Woodland. 2021 · 2021
Later among the works it cites.
End-to-end neural diarization: From transformer to conformer
Yi-Chieh Liu, Eunjung Han, Chul Lee, and Andreas Stolcke. 2021 · 2021
Later among the works it cites.
Online streaming end-to-end neural diarization handling overlapping speech and flexible numbers of speakers
Yawen Xue, Shota Horiguchi, Yusuke Fujita, Yuki Takashima, Shinji Watanabe, Leibny Paola García Perera, and Kenji Nagamatsu. 2021 · 2021
Later among the works it cites.
Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implementation and analysis on standard tasks
Federico Landini, Ján Profant, Mireia Diez, and Lukáš Burget. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mireia Díez, Lukás Burget, Shuai Wang, Johan Rohdin, and Jan Cernocký. 2019 · 2019
Cited alongside, same era.
LSTM based similarity measurement with spectral clustering for speaker diarization
Qingjian Lin, Ruiqing Yin, Ming Li, Hervé Bredin, and Claude Barras. 2019 · 2019
Cited alongside, same era.
The second DIHARD diarization challenge: Dataset, task, and baselines
Neville Ryant, Kenneth Church, Christopher Cieri, Alejandrina Cristià, Jun Du, Sriram Ganapathy, and Mark Liberman. 2019 · 2019
Cited alongside, same era.
Cn-celeb: A challenging chinese speaker recognition dataset
Y. Fan, J. W. Kang, L. T. Li, K. C. Li, H. L. Chen, S. T. Cheng, P. Y. Zhang, Z. Y. Zhou, Y. Q. Cai, and D. Wang. 2020 · 2020
Cited alongside, same era.
End-to-end speaker diarization for an unknown number of speakers with encoder-decoder based attractors
Shota Horiguchi, Yusuke Fujita, Shinji Watanabe, Yawen Xue, and Kenji Nagamatsu. 2020 · 2020
Cited alongside, same era.
Target-speaker voice activity detection: A novel approach for multi-speaker diarization in a dinner party scenario
Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach, and et al. 2020 · 2020
Cited alongside, same era.
Speaker turn modeling for dialogue act classification
Zihao He, Leili Tavabi, Kristina Lerman, and Mohammad Soleymani. 2021 · 2021
Cited alongside, same era.
Incorporating end-to-end framework into target-speaker voice activity detection
Weiqing Wang and Ming Li. 2022 · 2022
Closest in time.
Cross-channel attention-based target speaker voice activity detection: Experimental results for m2met challenge
Weiqing Wang, Xiaoyi Qin, and Ming Li. 2022 · 2022
Closest in time.
M2met: The icassp 2022 multi-channel multi-party meeting transcription challenge
Fan Yu, Shiliang Zhang, Yihui Fu, Lei Xie, Siqi Zheng, Zhihao Du, Weilong Huang, Pengcheng Guo, Zhijie Yan, Bin Ma, Xin Xu, and Hui Bu. 2022a · 2022
Closest in time.
Summary on the ICASSP 2022 multi-channel multi-party meeting transcription grand challenge
Fan Yu, Shiliang Zhang, Pengcheng Guo, Yihui Fu, Zhihao Du, Siqi Zheng, Weilong Huang, Lei Xie, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong Aik Lee, Zhijie Yan, Bin Ma, Xin Xu, and Hui Bu. 2022b · 2022
Closest in time.
Summary on the ICASSP 2022 multi-channel multi-party meeting transcription grand challenge
Fan Yu, Shiliang Zhang, Pengcheng Guo, Yihui Fu, Zhihao Du, Siqi Zheng, Weilong Huang, Lei Xie, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong Aik Lee, Zhijie Yan, Bin Ma, Xin Xu, and Hui Bu. 2022c · 2022
Closest in time.
Towards end-to-end speaker diarization with generalized neural speaker clustering
Chunlei Zhang, Jiatong Shi, Chao Weng, Meng Yu, and Dong Yu. 2022 · 2022
Closest in time.
The cuhk-tencent speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge
Naijun Zheng, Na Li, Xixin Wu, Lingwei Meng, Jiawen Kang, Haibin Wu, Chao Weng, Dan Su, and Helen Meng. 2022 · 2022
Closest in time.
Reformulating speaker diarization as community detection with emphasis on topological structure
Siqi Zheng and Hongbin Suo. 2022 · 2022
Closest in time.