Fetching the paper…
Reading the bibliography…
Attractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets.
“2000 NIST Speaker Recognition Evaluation,” https://catalog.ldc.upenn.edu/LDC2001S97
2000
Earlier work this paper cites.
“Constrained k-means clustering with background knowledge,”
Kiri Wagstaff, Claire Cardie, Seth Rogers, Stefan Schroedl, et al., · 2001
Earlier work this paper cites.
“An improved Cop-Kmeans clustering for solving constraint violation based on MapReduce framework,”
Yan Yang, Tonny Rutayisire, Chao Lin, Tianrui Li, and Fei Teng, · 2013
Earlier work this paper cites.
“LibriSpeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R. Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Semantic instance segmentation via deep metric learning,” arXiv:1703.10277, 2017
Alireza Fathi, Zbigniew Wojna, Vivek Rathod, Peng Wang, Hyun Oh Song, Sergio Guadarrama, and Kevin P Murphy, · 2017
Earlier work this paper cites.
“Deep attractor network for single-microphone speaker separation,”
Zhuo Chen, Yi Luo, and Nima Mesgarani, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Recurrent pixel embedding for instance grouping,”
Shu Kong and Charless C Fowlkes, · 2018
Earlier work this paper cites.
“Speaker diarization with LSTM,”
Quan Wang, Carlton Downey, Li Wan, Phlip Andrew Mansfield, and Ignacio Lopez Moreno, · 2018
Earlier work this paper cites.
“Alternative objective functions for deep clustering,”
Zhong-Qiu Wang, Johathan Le Roux, and John R Hershey, · 2018
Earlier work this paper cites.
“TasNet: time-domain audio separation network for real-time, single-channel speech separation,”
Yi Luo and Nima Mesgarani, · 2018
Earlier work this paper cites.
“End-to-end neural speaker diarization with self-attention,”
Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Yawen Xue, Kenji Nagamatsu, and Shinji Watanabe, · 2019
Earlier work this paper cites.
“Joint speech recognition and speaker diarization via sequence transduction,”
Laurent El Shafey, Hagen Soltau, and Izhak Shafran, · 2019
Earlier work this paper cites.
“Advances in online audio-visual meeting transcription,”
Takuya Yoshioka, Igor Abramovski, Cem Aksoylar, Zhuo Chen, Moshe David, Dimitrios Dimitriadis, Yifan Gong, Ilya Gurvich, Xuedong Huang, Yan Huang, Aviv Hurvitz, Li Jiang, Sharon Koubi, Eyal Krupka, Ido Leichter, Changliang Liu, Partha Parthasarathy, Alon Vinnikov, Lingfeng Wu, Xiong Xiao, Wayne Xiong, Huaming Wang, Zhenghao Wang, Jun Zhang, Yong Zhao, and Tianyan Zhou, · 2019
Earlier work this paper cites.
“The Second DIHARD Diarization Challenge: Dataset, task, and baselines,”
Neville Ryant, Kenneth Church, Christopher Cieri, Alejandrina Cristia, Jun Du, Sriram Ganapathy, and Mark Liberman, · 2019
Cited alongside, same era.
“Fully supervised speaker diarization,”
Aonan Zhang, Quan Wang, Zhenyao Zhu, John Paisley, and Chong Wang, · 2019
Cited alongside, same era.
“Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Cited alongside, same era.
“Recursive speech separation for unknown number of speakers,”
Naoya Takahashi, Sudarsanam Parthasaarathy, Nabarun Goswami, and Yuki Mitsufuji, · 2019
Cited alongside, same era.
“Target-speaker voice activity detection: a novel approach for multi-speaker diarization in a dinner party scenario,”
Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach, Yuri Khokhlov, Mariya Korenevskaya, Ivan Sorokin, Tatiana Timofeeva, Anton Mitrofanov, Andrei Andrusenko, Ivan Podluzhny, Aleksandr Laptev, and Aleksei Romanenko, · 2020
Cited alongside, same era.
“BUT system for the Second DIHARD Speech Diarization Challenge,”
Federico Landini, Shuai Wang, Mireia Diez, Lukáš Burget, Pavel Matějka, Kateřina Žmolíková, Ladislav Mošner, Anna Silnova, Oldřich Plchot, Ondřej Novotnỳ, Hossein Zeinali, and Johan Rohdin, · 2020
Later among the works it cites.
“A review of speaker diarization: Recent advances with deep learning,” arXiv:2101.09624, 2021
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu J. Han, Shinji Watanabe, and Shrikanth Narayanan, · 2021
Closest in time.
“The third DIHARD diarization challenge,”
Neville Ryant, Prachi Singh, Venkat Krishnamohan, Rajat Varma, Kenneth Church, Christopher Cieri, Jun Du, Sriram Ganapathy, and Mark Liberman, · 2021
Closest in time.
“End-to-end neural diarization: From transformer to conformer,”
Yi Chieh Liu, Eunjung Han, Chul Lee, and Andreas Stolcke, · 2021
Closest in time.
“End-to-end speaker diarization conditioned on speech activity and overlap detection,”
Yuki Takashima, Yusuke Fujita, Shinji Watanabe, Shota Horiguchi, Paola García, and Kenji Nagamatsu, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“CHiME-6 Challenge: Tackling multispeaker speech recognition for unsegmented recordings,”
Shinji Watanabe, Michael Mandel, Jon Barker, Emmanuel Vincent, Ashish Arora, Xuankai Chang, Sanjeev Khudanpur, Vimal Manohar, Daniel Povey, Desh Raj, David Snyder, Aswin Shanmugam Subramanian, Jan Trmal, Bar Ben Yair, Christoph Boeddeker, Zhaoheng Ni, Yusuke Fujita, Shota Horiguchi, Naoyuki Kanda, Takuya Yoshioka, and Neville Ryant, · 2020
Cited alongside, same era.
“End-to-end speaker diarization for an unknown number of speakers with encoder-decoder based attractors,”
Shota Horiguchi, Yusuke Fujita, Shinji Wananabe, Yawen Xue, and Kenji Nagamatsu, · 2020
Cited alongside, same era.
“Continuous speech separation: Dataset and analysis,”
Zhuo Chen, Takuya Yoshioka, Liang Lu, Tianyan Zhou, Zhong Meng, Yi Luo, Jian Wu, Xiong Xiao, and Jinyu Li, · 2020
Cited alongside, same era.
“Optimal mapping loss: A faster loss for end-to-end speaker diarization,”
Qingjian Lin, Tingle Li, Lin Yang, Junjie Wang, and Ming Li, · 2020
Cited alongside, same era.
“Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap,”
Tae Jin Park, Kyu J. Han, Manoj Kumar, and Shrikanth Narayanan, · 2020
Cited alongside, same era.
“Overlap-aware diarization: Resegmentation using neural end-to-end overlapped speech detection,”
Latané Bullock, Hervé Bredin, and Leibny Paola Garcia-Perera, · 2020
Cited alongside, same era.
“Personal VAD: Speaker-conditioned voice activity detection,”
Shaojin Ding, Quan Wang, Shuo-yiin Chang, Li Wan, and Ignacio Lopez Moreno, · 2020
Cited alongside, same era.
Closest in time.
“End-to-end diarization for variable number of speakers with local-global networks and discriminative speaker embeddings,”
Soumi Maiti, Hakan Erdogan, Kevin Wilson, Scott Wisdom, Shinji Watanabe, and John R. Hershey, · 2021
Closest in time.
“BW-EDA-EEND: Streaming end-to-end neural speaker diarization for a variable number of speakers,”
Eunjung Han, Chul Lee, and Andreas Stolcke, · 2021
Closest in time.
Shota Horiguchi, Yusuke Fujita, Shinji Watanabe, Yawen Xue, and Paola García, · 2021
Closest in time.
“Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds,”
Keisuke Kinoshita, Marc Delcroix, and Naohiro Tawara, · 2021
Closest in time.
“Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech,”
Keisuke Kinoshita, Marc Delcroix, and Naohiro Tawara, · 2021
Closest in time.
“ResNeXt and Res2Net structures for speaker verification,”
Tianyan Zhou, Yong Zhao, and Jian Wu, · 2021
Closest in time.
“Discriminative neural clustering for speaker diarisation,”
Qiujia Li, Florian L. Kreyssig, Chao Zhang, and Philip C. Woodland, · 2021
Closest in time.
“End-to-end speaker diarization as post-processing,”
Shota Horiguchi, Paola Garcia, Yusuke Fujita, Shinji Watanabe, and Kenji Nagamatsu, · 2021
Closest in time.
“Wavesplit: End-to-end speech separation by speaker clustering,”
Neil Zeghidour and David Grangier, · 2021
Closest in time.
“The Hitachi-JHU DIHARD III system: Competitive end-to-end neural diarization and x-vector clustering systems combined by DOVER-Lap,”
Shota Horiguchi, Nelson Yalta, Paola Garcia, Yuki Takashima, Yawen Xue, Desh Raj, Zili Huang, Yusuke Fujita, Shinji Watanabe, and Sanjeev Khudanpur, · 2021
Closest in time.
“Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: Theory, implementation and analysis on standard tasks,”
Federico Landini, Ján Profant, Mireia Diez, and Lukáš Burget, · 2022
Closest in time.