Fetching the paper…
Reading the bibliography…
Speaker diarization relies on the assumption that speech segments corresponding to a particular speaker are concentrated in a specific region of the speaker space; a region which represents that speaker's identity.
The linguistic individual: Self-expression in language and linguistics
Barbara Johnstone, · 1996
Earlier work this paper cites.
“Maximum Likelihood Linear Transformations for HMM-based Speech Recognition,”
Mark JF Gales, · 1998
Earlier work this paper cites.
“Speaker, Environment and Channel Change Detection and Clustering via the Bayesian Information Criterion,”
Scott Chen and Ponani Gopalakrishnan, · 1998
Earlier work this paper cites.
“The Rules Behind Roles: Identifying Speaker Role in Radio Broadcasts,”
Regina Barzilay, Michael Collins, Julia Hirschberg, and Steve Whittaker, · 2000
Earlier work this paper cites.
“SRILM–An Extensible Language Modeling Toolkit,”
Andreas Stolcke, · 2002
Earlier work this paper cites.
“The Fisher Corpus: a Resource for the Next Generations of Speech-to-Text,”
Christopher Cieri, David Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“Probabilistic Linear Discriminant Analysis,”
Sergey Ioffe, · 2006
Earlier work this paper cites.
“The IBM Rich Transcription 2007 Speech-to-Text Systems for Lecture Meetings,”
Jing Huang, Etienne Marcheret, Karthik Visweswariah, Vit Libal, and Gerasimos Potamianos, · 2007
Earlier work this paper cites.
“Probabilistic Linear Discriminant Analysis for Inferences about Identity,”
Simon JD Prince and James H Elder, · 2007
Earlier work this paper cites.
“Exploiting Intra-conversation Variability for Speaker Diarization,”
Stephen Shum, Najim Dehak, Ekapol Chuangsuwanich, Douglas Reynolds, and James Glass, · 2011
Earlier work this paper cites.
“Automatic Role Recognition in Multiparty Conversations: An Approach Based on Turn Organization, Prosody, and Conditional Random Fields,”
Hugues Salamin and Alessandro Vinciarelli, · 2011
Earlier work this paper cites.
“Speaker Role Recognition Using Question Detection and Characterization,”
Thierry Bazillon, Benjamin Maza, Michael Rouvier, Frederic Bechet, and Alexis Nasr, · 2011
Earlier work this paper cites.
“Robust Speaker Turn Role Labeling of TV Broadcast News Shows,”
Géraldine Damnati and Delphine Charlet, · 2011
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“Speaker Diarization: A Review of Recent Research,”
Xavier Anguera, Simon Bozonnet, Nicholas Evans, Corinne Fredouille, Gerald Friedland, and Oriol Vinyals, · 2012
Earlier work this paper cites.
“Incorporation of the ASR Output in Speaker Segmentation and Clustering within the Task of Speaker Diarization of Broadcast Streams,”
Jan Silovsky, Jindrich Zdansky, Jan Nouza, Petr Cerva, and Jan Prazak, · 2012
Earlier work this paper cites.
“Speaker Adaptation of Neural Network Acoustic Models Using i-vectors,”
George Saon, Hagen Soltau, David Nahamoo, and Michael Picheny, · 2013
Earlier work this paper cites.
“Speaker Role Recognition on TV Broadcast Documents,”
Benjamin Bigot, Corinne Fredouille, and Delphine Charlet, · 2013
Earlier work this paper cites.
“Distributed Representations of Words and Phrases and Their Compositionality,”
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean, · 2013
Cited alongside, same era.
“Speaker Diarization with PLDA i-vector Scoring and Unsupervised Calibration,”
Gregory Sell and Daniel Garcia-Romero, · 2014
Cited alongside, same era.
“Scaling Up the Evaluation of Psychotherapy: Evaluating Motivational Interviewing Fidelity via Statistical Text Classification,”
David C Atkins, Mark Steyvers, Zac E Imel, and Padhraic Smyth, · 2014
Cited alongside, same era.
“Multimodal Embedding Fusion for Robust Speaker Role Recognition in Video Broadcast,”
Michael Rouvier, Sebastien Delecraz, Benoit Favre, Meriem Bendris, and Frederic Bechet, · 2015
Cited alongside, same era.
“A Technology Prototype System for Rating Therapist Empathy from Audio Recordings in Addiction Counseling,”
Bo Xiao, Chewei Huang, Zac E Imel, David C Atkins, Panayiotis Georgiou, and Shrikanth S Narayanan, · 2016
Cited alongside, same era.
“Recurrent Neural Network Based Speaker Change Detection from Text Transcription Applied in Telephone Speaker Diarization System,”
Zbyněk Zajíc, Daniel Soutner, Marek Hrúz, Luděk Müller, and Vlasta Radová, · 2018
Later among the works it cites.
“Combined Speaker Clustering and Role Recognition in Conversational Speech,”
Nikolaos Flemotomos, Pavlos Papadopoulos, James Gibson, and Shrikanth Narayanan, · 2018
Later among the works it cites.
“Role Specific Lattice Rescoring for Speaker Role Recognition from Speech Recognition Outputs,”
Nikolaos Flemotomos, Panayiotis Georgiou, and Shrikanth Narayanan, · 2018
Later among the works it cites.
“Role annotated speech recognition for conversational interactions,”
Nikolaos Flemotomos, Zhuohao Chen, David C Atkins, and Shrikanth Narayanan, · 2018
Later among the works it cites.
“X-vectors: Robust DNN Embeddings for Speaker Recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Investigation of Segmentation in i-vector Based Speaker Diarization of Telephone Speech,”
Zbyněk Zajíc, Marie Kunešová, and Vlasta Radová, · 2016
Cited alongside, same era.
“End-to-End Sequence Labeling via Bi-directional LSTM-CNNs-CRF,”
Xuezhe Ma and Eduard Hovy, · 2016
Cited alongside, same era.
“Dependency Based Embeddings for Sentence Classification Tasks,”
Alexandros Komninos and Suresh Manandhar, · 2016
Cited alongside, same era.
“Hierarchical RNN with Static Sentence-Level Attention for Text-Based Speaker Change Detection,”
Zhao Meng, Lili Mou, and Zhi Jin, · 2017
Cited alongside, same era.
“LSTM Neural Network-Based Speaker Segmentation Using Acoustic and Language Modelling,”
Miquel Àngel India Massana, José Adrián Rodríguez Fonollosa, and Francisco Javier Hernando Pericás, · 2017
Cited alongside, same era.
“Speaker Diarization Using Deep Neural Network Embeddings,”
Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, and Alan McCree, · 2017
Cited alongside, same era.
“Tristounet: Triplet Loss for Speaker Turn Embedding,”
Hervé Bredin, · 2017
Cited alongside, same era.
“Speaker Diarization with LSTM,”
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno, · 2018
Later among the works it cites.
“Neural Speech Turn Segmentation and Affinity Propagation for Speaker Diarization,”
Ruiqing Yin, Hervé Bredin, and Claude Barras, · 2018
Later among the works it cites.
“NCRF++: An Open-source Neural Sequence Labeling Toolkit,”
Jie Yang and Yue Zhang, · 2018
Later among the works it cites.
“A Systematic Comparison of contemporary Automatic Speech Recognition Engines for Conversational Clinical Speech,”
Jodi Kodish-Wachs, Emin Agassi, Patrick Kenny III, and J Marc Overhage, · 2018
Later among the works it cites.
“Speaker Diarization with Lexical Information,”
Tae Jin Park, Kyu J. Han, Jing Huang, Xiaodong He, Bowen Zhou, Panayiotis Georgiou, and Shrikanth Narayanan, · 2019
Closest in time.
“Joint Speech Recognition and Speaker Diarization via Sequence Transduction,”
Laurent El Shafey, Hagen Soltau, and Izhak Shafran, · 2019
Closest in time.
“Meeting Transcription Using Asynchronous Distant Microphones,”
Takuya Yoshioka, Dimitrios Dimitriadis, Andreas Stolcke, William Hinthorn, Zhuo Chen, Michael Zeng, and Xuedong Huang, · 2019
Closest in time.
“The Second DIHARD Challenge: System Description for USC-SAIL Team,”
Tae Jin Park, Manoj Kumar, Nikolaos Flemotomos, Monisankha Pal, Raghuveer Peri, Rimita Lahiri, Panayiotis Georgiou, and Shrikanth Narayanan, · 2019
Closest in time.
“Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network,”
Monisankha Pal, Manoj Kumar, Raghuveer Peri, Tae Jin Park, So Hyun Kim, Catherine Lord, Somer Bishop, and Shrikanth Narayanan, · 2019
Closest in time.
“Fully Supervised Speaker Diarization,”
Aonan Zhang, Quan Wang, Zhenyao Zhu, John Paisley, and Chong Wang, · 2019
Closest in time.
“End-to-End Neural Speaker Diarization with Self-Attention,”
Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Yawen Xue, Kenji Nagamatsu, and Shinji Watanabe, · 2019
Closest in time.
“Positioning Oneself in Different roles: Structural and Lexical Measures of Power Relations Between Speakers in Map Task Corpus,”
Vered Silber-Varod, Sarit Malayev, and Anat Lerner, · 2020
Closest in time.