Fetching the paper…
Reading the bibliography…
Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task.
Y. Ito, “Representation of functions by superpositions of a step or sigmoid function and their applications to neural network theory,”
1991
Earlier work this paper cites.
D. M. David Graff, Kevin Walker, “Switchboard cellular part 1 audio,” https://catalog.ldc.upenn.edu/LDC2001S13, 2001
2001
Earlier work this paper cites.
M. Rimer and T. Martinez, “Softprop: softmax neural network backpropagation learning,” in
2004
Earlier work this paper cites.
J.-M. Cheng and H.-C. Wang, “A method of estimating the equal error rate for automatic speaker verification,” in
2004
Earlier work this paper cites.
Z.-L. Zhang and M.-L. Zhang, “Multi-instance multi-label learning with application to scene classification,” in
2007
Earlier work this paper cites.
V. Tiwari, “Mfcc and its applications in speaker recognition,”
2010
Earlier work this paper cites.
K. P. Murphy,
2012
Earlier work this paper cites.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in
2014
Earlier work this paper cites.
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” in
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Cited alongside, same era.
Z. Yang, D. Yang, C. Dyer, X. He, A. Smola, and E. Hovy, “Hierarchical attention networks for document classification,” in
2016
Cited alongside, same era.
Y. Xu, Q. Huang, W. Wang, P. Foster, S. Sigtia, P. J. Jackson, and M. D. Plumbley, “Unsupervised feature learning based on deep models for environmental audio tagging,”
2017
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,”
2017
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in
2018
Later among the works it cites.
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, “Self-attentive speaker embeddings for text-independent speaker verification.” in
2018
Later among the works it cites.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,”
2018
Later among the works it cites.
F. R. rahman Chowdhury, Q. Wang, I. L. Moreno, and L. Wan, “Attention-based models for text-dependent speaker verification,” in
2018
Later among the works it cites.
Y. Jia, M. Johnson, W. Macherey, R. J. Weiss, Y. Cao, C.-C. Chiu, N. Ari, S. Laurenzo, and Y. Wu, “Leveraging weakly supervised data to improve end-to-end speech-to-text translation,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Q. Wang, K. Okabe, K. A. Lee, H. Yamamoto, and T. Koshinaka, “Attention mechanism in speaker recognition: What does it learn in deep speaker embedding?” in
2018
Cited alongside, same era.
M. Karu and T. Alumäe, “Weakly supervised training of speaker identification models,” in
2018
Cited alongside, same era.
Z.-H. Zhou, “A brief introduction to weakly supervised learning,”
2018
Cited alongside, same era.
Y. Xu, Q. Kong, W. Wang, and M. D. Plumbley, “Large-scale weakly supervised audio classification using gated convolutional neural network,” in
2018
Cited alongside, same era.
W. Liu, R. Qin, and F. Su, “Weakly supervised classification of time-series of very high resolution remote sensing images by transfer learning,”
2019
Later among the works it cites.
X. Xu, G. Li, G. Xie, J. Ren, and X. Xie, “Weakly supervised deep semantic segmentation using cnn and elm with semantic candidate regions,”
2019
Later among the works it cites.
Y. Pan, B. Mirheidari, M. Reuber, A. Venneri, D. Blackburn, and H. Christensen, “Automatic hierarchical attention neural network for detecting ad,” in
2019
Later among the works it cites.
Y. Shi, Q. Huang, and T. Hain, “H-vectors: Utterance-level speaker embedding using a hierarchical attention model,” in
2020
Closest in time.