Fetching the paper…
Reading the bibliography…
Speech applications dealing with conversations require not only recognizing the spoken words, but also determining who spoke when.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,”
1995
Earlier work this paper cites.
T. Robinson, M. Hochberg, and S. Renals, “The use of recurrent neural networks in continuous speech recognition,” in
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
J. Ajmera and C. Wooters, “A robust speaker clustering algorithm,” in
2003
Earlier work this paper cites.
Anonymous, “The Rich Transcription Fall 2003 (RT-03F) Evaluation Plan,” NIST, Tech. Rep., 2003
2003
Earlier work this paper cites.
L. Canseco-Rodriguez, L. Lamel, and J.-L. Gauvain, “Speaker diarization from speech transcripts,” in
2004
Earlier work this paper cites.
S. E. Tranter and D. A. Reynolds, “An overview of automatic speaker diarization systems,”
2006
Earlier work this paper cites.
C. Barras, X. Zhu, S. Meignier, and J.-L. Gauvain, “Multistage speaker diarization of broadcast news,”
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, “Speaker diarization: A review of recent research,”
2012
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
A. Graves, A. Mohamed, and G. E. Hinton, “Speech recognition with deep recurrent neural networks,” in
2013
Cited alongside, same era.
S. Virpioja, P. Smit, S.-A. Grönroos, and M. Kurimo, “Morfessor 2.0: Python implementation and extensions for morfessor baseline,” Aalto University, Tech. Rep., 2013
2013
Cited alongside, same era.
G. Sell and D. Garcia-Romero, “Speaker diarization with PLDA i-vector scoring and unsupervised calibration,” in
2014
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
R. Zazo, T. N. Sainath, G. Simko, and C. Parada, “Feature learning with raw-waveform cldnns for voice activity detection,” in
2016
Cited alongside, same era.
H. Soltau, H. Liao, and H. Sak, “Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition,” in
2017
Later among the works it cites.
H. Soltau, H. Liao, and H. Sak, “Reducing the computational complexity for whole word models,” in
2017
Later among the works it cites.
Q. Wang, C. Downey, L. Wan, P. A. Mansfield, and I. L. Moreno, “Speaker diarization with LSTM,” in
2018
Later among the works it cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in
2018
Later among the works it cites.
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, “Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,” in
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Heigold, I. Moreno, S. Bengio, and N. M. Shazeer, “End-to-end text-dependent speaker verification,” in
2016
Cited alongside, same era.
D. Garcia-Romero, D. Snyder, G. Sell, D. Povey, and A. McCree, “Speaker diarization using deep neural network embeddings,” in
2017
Cited alongside, same era.
H. Bredin, “Tristounet: Triplet loss for speaker turn embedding,” in
2017
Cited alongside, same era.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers
2017
Cited alongside, same era.
K. C. Sim, A. Narayanan, T. Bagby, T. N. Sainath, and M. Bacchiani, “Improving the efficiency of forward-backward algorithm using batched computation in tensorflow,” in
2017
Cited alongside, same era.
Later among the works it cites.
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, “Fully supervised speaker diarization,”
2018
Later among the works it cites.
T. J. Park and P. G. Georgiou, “Multimodal speaker segmentation and diarization using lexical and acoustic cues via sequence to sequence neural networks,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Bagby and K. Rao, “Efficient implementation of recurrent neural network transducer in tensorflow,” in
2018
Later among the works it cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in
2018
Later among the works it cites.