Fetching the paper…
Reading the bibliography…
While there has been substantial amount of work in speaker diarization recently, there are few efforts in jointly employing lexical and acoustic information for speaker segmentation.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,”
1989
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
J. J. Godfrey and E. Holliman, “Switchboard-1 release 2,”
1997
Earlier work this paper cites.
S. Chen, P. Gopalakrishnan
1998
Earlier work this paper cites.
A. Tritschler and R. A. Gopinath, “Improved speaker segmentation and segments clustering using the bayesian information criterion,” in
1999
Earlier work this paper cites.
L. Canseco-Rodriguez, L. Lamel, and J.-L. Gauvain, “Speaker diarization from speech transcripts.” ICSLP, 2004
2004
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “Fisher english training speech parts 1 and 2,”
2004
Earlier work this paper cites.
J. G. Fiscus, J. Ajot, M. Michel, and J. S. Garofolo, “The rich transcription 2006 spring meeting recognition evaluation,” in
2006
Earlier work this paper cites.
Y. Esteve, S. Meignier, P. Deléglise, and J. Mauclair, “Extracting true speaker identities from transcriptions,” in
2007
Earlier work this paper cites.
P. G. Georgiou, M. P. Black, and S. S. Narayanan, “Behavioral signal processing for understanding (distressed) dyadic interactions: some recent developments,” in
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Cited alongside, same era.
S. Narayanan and P. G. Georgiou, “Behavioral signal processing: Deriving human behavioral informatics from speech and language,”
2013
Cited alongside, same era.
M. Rouvier, G. Dupuy, P. Gay, E. Khoury, T. Merlin, and S. Meignier, “An open-source state-of-the-art toolbox for broadcast news diarization,” in
2013
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in
2014
Cited alongside, same era.
2014
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Later among the works it cites.
R. Nallapati, B. Zhou, C. Gulcehre, B. Xiang
2016
Later among the works it cites.
R. Yin, H. Bredin, and C. Barras, “Speaker change detection in broadcast tv using bidirectional long short-term memory networks,” in
2017
Later among the works it cites.
Q. Wang, C. Downey, L. Wan, P. A. Mansfield, and I. L. Moreno, “Speaker diarization with lstm,”
2017
Later among the works it cites.
D. Garcia-Romero, D. Snyder, G. Sell, D. Povey, and A. McCree, “Speaker diarization using deep neural network embeddings,” in
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
B. Desplanques, K. Demuynck, and J.-P. Martens, “Factor analysis for speaker segmentation and improved speaker diarization,” in
2015
Cited alongside, same era.
B. Xiao, P. Georgiou, Z. E. Imel, D. Atkins, and S. Narayanan, ““Rate my therapist”: Automated detection of empathy in drug and alcohol counseling via speech and language processing,”
2015
Cited alongside, same era.
B. Xiao, C. Huang, Z. E. Imel, D. C. Atkins, P. Georgiou, and S. S. Narayanan, “A technology prototype system for rating therapist empathy from audio recordings in addiction counseling,”
2016
Cited alongside, same era.
Later among the works it cites.
A. Jati and P. Georgiou, “Speaker2vec: Unsupervised learning and adaptation of a speaker manifold using deep neural networks with an evaluation on speaker segmentation,”
2017
Later among the works it cites.
M. À. India Massana, J. A. Rodríguez Fonollosa, and F. J. Hernando Pericás, “Lstm neural network-based speaker segmentation using acoustic and language modelling,” in
2017
Later among the works it cites.
——, “Neural predictive coding using convolutional neural networks towards unsupervised learning of speaker characteristics,”
2018
Closest in time.
J. Lyons, “Python speech features,” https://github.com/jameslyons/python-speech-features, 2017, accessed: 2018-03-23
2018
Closest in time.