Fetching the paper…
Reading the bibliography…
Several advances have been made recently towards handling overlapping speech for speaker diarization.
“The Hungarian method for the assignment problem,”
Harold W. Kuhn, · 1955
Earlier work this paper cites.
“Maximum bounded 3-dimensional matching is MAX SNP-complete,”
V. Kann, · 1991
Earlier work this paper cites.
“A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (ROVER),”
Jonathan G. Fiscus, · 1997
Earlier work this paper cites.
“Complexity and approximation: combinatorial optimization problems and their approximability properties,”
G. Ausiello, A. Marchetti-Spaccamela, P. Crescenzi, G. Gambosi, M. Protasi, and V. Kann, · 1999
Earlier work this paper cites.
“Posterior probability decoding, confidence estimation and system combination,”
Gunnar Evermann and PC Woodland, · 2000
Earlier work this paper cites.
“The SRI March 2000 Hub-5 conversational speech transcription system,”
Andreas Stolcke, Harry Bratt, John Butzberger, Horacio Franco, V. R. Rao Gadde, Madelaine Plauché, Colleen Richey, Elizabeth Shriberg, Kemal Sönmez, Fuliang Weng, and Jing Zheng, · 2000
Earlier work this paper cites.
“The AMI meeting corpus: A pre-announcement,”
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Maël Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska Masson, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner, · 2005
Earlier work this paper cites.
“An overview of automatic speaker diarization systems,”
Sue Tranter and Douglas A. Reynolds, · 2006
Earlier work this paper cites.
“Incremental assignment problem,”
I. H. Toroslu and G. Üçoluk, · 2007
Earlier work this paper cites.
“The AMI system for the transcription of speech in meetings,”
Thomas Hain, Vincent Wan, Lukás Burget, Martin Karafiát, John Dines, Jithendra Vepa, Giulia Garau, and Mike Lincoln, · 2007
Earlier work this paper cites.
“Acoustic beamforming for speaker diarization of meetings,”
Xavier Anguera Miró, Chuck Wooters, and Javier Hernando, · 2007
Earlier work this paper cites.
“Overlapped speech detection for improved speaker diarization in multiparty meetings,”
Kofi Boakye, Beatriz Trueba-Hornero, Oriol Vinyals, and Gerald Friedland, · 2008
Earlier work this paper cites.
“Speech overlap detection in a two-pass speaker diarization system,”
Marijn Huijbregts, David A. van Leeuwen, and Franciska de Jong, · 2009
Earlier work this paper cites.
“Ensemble-based classifiers,”
Lior Rokach, · 2009
Earlier work this paper cites.
“Reducibility among combinatorial problems,”
R. Karp, · 2010
Earlier work this paper cites.
“Speech dereverberation based on variance-normalized delayed linear prediction,”
T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, and B. Juang, · 2010
Cited alongside, same era.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2011
Cited alongside, same era.
“Speaker diarization: A review of recent research,”
Xavier Anguera Miró, Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille, Gerald Friedland, and Oriol Vinyals, · 2012
Cited alongside, same era.
“Speaker diarization of overlapping speech based on silence distribution in meeting recordings,”
Sree Harsha Yella and Fabio Valente, · 2012
Cited alongside, same era.
“Detecting overlapping speech with long short-term memory recurrent neural networks,”
Jürgen T. Geiger, Florian Eyben, Björn W. Schuller, and Gerhard Rigoll, · 2013
Cited alongside, same era.
“NARA-WPE: A Python package for weighted prediction error dereverberation in Numpy and Tensorflow for online and offline processing,”
Lukas Drude, J. Heymann, Christoph Böddeker, and R. Haeb-Umbach, · 2018
Later among the works it cites.
“Detection of overlapping speech for the purposes of speaker diarization,”
Marie Kunesová, Marek Hrúz, Zbynek Zajíc, and Vlasta Radová, · 2019
Later among the works it cites.
“Overlap-aware diarization: resegmentation using neural end-to-end overlapped speech detection,”
Latané Bullock, Hervé Bredin, and L. Paola García-Perera, · 2019
Later among the works it cites.
“DOVER: A method for combining diarization outputs,”
Andreas Stolcke and Takuya Yoshioka, · 2019
Later among the works it cites.
“Speech processing for digital home assistants: Combining signal processing with deep-learning techniques,”
Reinhold Haeb-Umbach, Shinji Watanabe, Tomohiro Nakatani, Michiel Bacchiani, Björn Hoffmeister, Michael L. Seltzer, Heiga Zen, and Mehrez Souden, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez-Moreno, and Javier Gonzalez-Dominguez, · 2014
Cited alongside, same era.
“Diarization resegmentation in the factor analysis subspace,”
Gregory Sell and Daniel Garcia-Romero, · 2015
Cited alongside, same era.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“Speaker diarization using deep neural network embeddings,”
Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, and Alan McCree, · 2017
Cited alongside, same era.
“Detecting overlapped speech on short timeframes using deep learning,”
Valentin Andrei, Horia Cucu, and Corneliu Burileanu, · 2017
Cited alongside, same era.
“Enhancing LSTM RNN-based speech overlap detection by artificially mixed data,”
Gerhard Hagerer, Vedhas Pandit, Florian Eyben, and Björn W. Schuller, · 2017
Cited alongside, same era.
“Speaker diarization with enhancing speech for the first DIHARD challenge,”
Lei Sun, Jun Du, Chao Jiang, Xueyang Zhang, Shan He, Bing Yin, and Chin-Hui Lee, · 2018
Cited alongside, same era.
Later among the works it cites.
“Meeting transcription using asynchronous distant microphones,”
Takuya Yoshioka, Dimitrios Dimitriadis, Andreas Stolcke, William Hinthorn, Zhuo Chen, Michael Zeng, and Xuedong Huang, · 2019
Later among the works it cites.
Yusuke Fujita, Shinji Watanabe, Shota Horiguchi, Yawen Xue, and Kenji Nagamatsu, · 2020
Closest in time.
“Speaker diarization with region proposal network,”
Zili Huang, Shinji Watanabe, Yusuke Fujita, Paola García, Yiwen Shao, Daniel Povey, and Sanjeev Khudanpur, · 2020
Closest in time.
Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach, Yuri Y. Khokhlov, Mariya Korenevskaya, Ivan Sorokin, Tatiana V. Timofeeva, Anton Mitrofanov, Andrei Andrusenko, Ivan Podluzhny, Aleksandr Laptev, and Aleksei Romanenko, · 2020
Closest in time.
“The JHU multi-microphone multi-speaker ASR system for the CHiME-6 challenge,”
Ashish Arora, Desh Raj, Aswin Shanmugam Subramanian, Ke Li, Bar Ben-Yair, Matthew Maciejewski, Piotr Zelasko, Paola Garcia, Shinji Watanabe, and Sanjeev Khudanpur, · 2020
Closest in time.
“CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,”
Shinji Watanabe, Michael Mandel, Jon Barker, and Emmanuel Vincent, · 2020
Closest in time.
“Microsoft speaker diarization system for the VoxCeleb speaker recognition challenge 2020,”
Xiong Xiao, Naoyuki Kanda, Z. Chen, Tianyan Zhou, T. Yoshioka, Sanyuan Chen, Yong Zhao, Gang Liu, Y. Wu, J. Wu, Shujie Liu, Jinyu Li, and Yifan Gong, · 2020
Closest in time.
“Continuous speech separation: Dataset and analysis,”
Zhuo Chen, Takuya Yoshioka, Liang Lu, Tianyan Zhou, Zhong Meng, Yi Luo, J. Wu, and Jinyu Li, · 2020
Closest in time.
“Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap,”
Tae Jin Park, Kyu J. Han, Manoj Kumar, and Shrikanth S. Narayanan, · 2020
Closest in time.
“Multi-class spectral clustering with overlaps for speaker diarization,”
Desh Raj, Zili Huang, and Sanjeev Khudanpur, · 2021
Closest in time.