Fetching the paper…
Reading the bibliography…
In this paper, we propose Discriminative Neural Clustering (DNC) that formulates data clustering with a maximum number of clusters as a supervised sequence-to-sequence learning problem.
“The subgroup algorithm for generating uniform random variables,”
P. Diaconis & M. Shahshahani, · 1987
Earlier work this paper cites.
“Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
R.J. Williams, · 1992
Earlier work this paper cites.
“Learning and development in neural networks: The importance of starting small,”
J.L. Elman, · 1993
Earlier work this paper cites.
“The AMI meeting corpus: A pre-announcement,”
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, I. McCowan, W. Post, D. Reidsma, & P. Wellner, · 2005
Earlier work this paper cites.
“An overview of automatic speaker diarization systems,”
S.E. Tranter & D.A. Reynolds, · 2006
Earlier work this paper cites.
“A spectral clustering approach to speaker diarization,”
H. Ning, M. Liu, H. Tang, & T.S. Huang, · 2006
Earlier work this paper cites.
“Acoustic beamforming for speaker diarization of meetings,”
X. Anguera, C. Wooters, & J. Hernando, · 2007
Earlier work this paper cites.
“Visualizing data using t-SNE,”
L.v.d. Maaten & G. Hinton, · 2008
Earlier work this paper cites.
“Curriculum learning,”
Y. Bengio, J. Louradour, R. Collobert, & J. Weston, · 2009
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P.J. Kenny, R. Dehak, P. Dumouchel, & P. Ouellet, · 2011
Earlier work this paper cites.
“Speaker diarization: A review of recent research,”
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, & O. Vinyals, · 2012
Earlier work this paper cites.
“Unsupervised methods for speaker diarization: An integrated and iterative approach,”
S.H. Shum, N. Dehak, R. Dehak, & J.R. Glass, · 2013
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, & Y. Bengio, · 2014
Earlier work this paper cites.
“Speaker diarisation and longitudinal linking in multi-genre broadcast data,”
P. Karanasou, M.J.F. Gales, P. Lanchantin, X. Liu, Y. Qian, L. Wang, P.C. Woodland, & C. Zhang, · 2015
Earlier work this paper cites.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
V. Peddinti, D. Povey, & S. Khudanpur, · 2015
Earlier work this paper cites.
The HTK Book
S. Young, G. Evermann, M. Gales, T. Hain, D. Kershaw, X. Liu, G. Moore, J. Odell, D. Ollason, D. Povey, A. Ragni, V. Valtchev, P. Woodland, & C. Zhang, · 2015
Cited alongside, same era.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
J.R. Hershey, Z. Chen, J. Le Roux, & S. Watanabe, · 2016
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, & Y. Bengio, · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q.V. Le, & O. Vinyals, · 2016
Cited alongside, same era.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, & J. Sun, · 2016
Cited alongside, same era.
“Layer normalization,”
J.L. Ba, J.R. Kiros, & G.E. Hinton, · 2016
Cited alongside, same era.
“Speaker diarization with LSTM,”
Q. Wang, C. Downey, L. Wan, P.A. Mansfield, & I.L. Moreno, · 2018
Later among the works it cites.
“Robust and discriminative speaker embedding via intra-class distance variance regularization,”
N. Le & J.M. Odobez, · 2018
Later among the works it cites.
“SpectralNet: Spectral Clustering using Deep Neural Networks,”
U. Shaham, K. Stanton, H. Li, B. Nadler, R. Basri, & Y. Kluger, · 2018
Later among the works it cites.
“Improved TDNNs using deep kernels and frequency dependent grid-RNNS,”
F.L. Kreyssig, C. Zhang, & P.C. Woodland, · 2018
Later among the works it cites.
“ESPnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Yalta, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, & T. Ochiai, · 2018
Later among the works it cites.
“Minimum Word Error Rate Training for Attention-Based Sequence-to-Sequence Models,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deep neural network embeddings for text-independent speaker verification,”
D. Snyder, D. Garcia-Romero, D. Povey, & S. Khudanpur, · 2017
Cited alongside, same era.
“Speaker diarization using deep neural network embeddings,”
D. Garcia-Romero, D. Snyder, G. Sell, D. Povey, & A. McCree, · 2017
Cited alongside, same era.
“Developing on-line speaker diarization system,”
D. Dimitriadis & P. Fousek, · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, & I. Polosukhin, · 2017
Cited alongside, same era.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
D. Yu, M. Kolbæk, Z.H. Tan, & J. Jensen, · 2017
Cited alongside, same era.
“SphereFace: Deep hypersphere embedding for face recognition,”
W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, & L. Song, · 2017
Cited alongside, same era.
R. Prabhavalkar, T. Sainath, Y. Wu, P. Nguyen, Z. Chen, C.C. Chiu, & A. Kannan, · 2018
Later among the works it cites.
“Speaker diarisation using 2D self-attentive combination of embeddings,”
G. Sun, C. Zhang, & P.C. Woodland, · 2019
Closest in time.
“LSTM based similarity measurement with spectral clustering for speaker diarization,”
Q. Lin, R. Yin, M. Li, H. Bredin, & C. Barras, · 2019
Closest in time.
“Fully supervised speaker diarization,”
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, & C. Wang, · 2019
Closest in time.
“End-to-end neural speaker diarization with self-attention,”
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, & S. Watanabe, · 2019
Closest in time.
“End-to-end neural speaker diarization with permutation-free objectives,”
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, & S. Watanabe, · 2019
Closest in time.
“PyHTK: Python library and ASR pipelines for HTK,”
C. Zhang, F.L. Kreyssig, Q. Li, & P.C. Woodland, · 2019
Closest in time.
“Monotonic infinite lookback attention for simultaneous machine translation,”
N. Arivazhagan, C. Cherry, W. Macherey, C.C. Chiu, S. Yavuz, R. Pang, W. Li, & C. Raffel, · 2019
Closest in time.
“Speaker diarization with session-level speaker embedding refinement using graph neural networks,”
J. Wang, X. Xiao, J. Wu, R. Ramamurthy, F. Rudzicz, & M. Brudno, · 2020
Closest in time.
“Monotonic multihead attention,”
X. Ma, J. Pino, J. Cross, L. Puzon, & J. Gu, · 2020
Closest in time.