Fetching the paper…
Reading the bibliography…
We introduce "Unspeech" embeddings, which are based on unsupervised learning of context feature representations for spoken language.
L. Hubert and P. Arabie, “Comparing partitions,”
1985
Earlier work this paper cites.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,” in
1990
Earlier work this paper cites.
J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah, “Signature verification using a ”Siamese” time delay neural network,” in
1994
Earlier work this paper cites.
H. Jin, F. Kubala, and R. Schwartz, “Automatic speaker clustering,” in
1997
Earlier work this paper cites.
B. Zhou and J. H. Hansen, “Unsupervised audio stream segmentation and clustering via the bayesian information criterion,” in
2000
Earlier work this paper cites.
A. Strehl, “Relationship-based clustering and cluster ensembles for high-dimensional data mining,” Ph.D. dissertation, University Of Texas at Austin, 2002
2002
Earlier work this paper cites.
M. Gutmann and A. Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” in
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Earlier work this paper cites.
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, “Speaker adaptation of neural network acoustic models using i-vectors.” in
2013
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean, “Distributed Representations of Words and Phrases and their Compositionality,” in
2013
Earlier work this paper cites.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in
2013
Earlier work this paper cites.
S. Bengio and G. Heigold, “Word embeddings for speech recognition,” in
2014
Earlier work this paper cites.
G. Synnaeve and E. Dupoux, “Weakly supervised multi-embeddings learning of acoustic models,”
2014
Cited alongside, same era.
A. Senior and I. Lopez-Moreno, “Improving DNN speaker independence with i-vector inputs,” in
2014
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2014
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,”
2014
Cited alongside, same era.
A. Rousseau, P. Deléglise, and Y. Estève, “Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks,” in
2014
Cited alongside, same era.
N. Zeghidour, G. Synnaeve, N. Usunier, and E. Dupoux, “Joint learning of speaker and phonetic similarities with siamese networks,” in
2016
Later among the works it cites.
K. Veselỳ, S. Watanabe, K. Žmolíková, M. Karafiát, L. Burget, and J. H. Černockỳ, “Sequence summarizing neural network for speaker adaptation,” in
2016
Later among the works it cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI.” in
2016
Later among the works it cites.
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, “Deep neural network embeddings for text-independent speaker verification,” in
2017
Later among the works it cites.
L. Li, Y. Chen, Y. Shi, Z. Tang, and D. Wang, “Deep speaker feature learning for text-independent speaker verification,” in
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in
2015
Cited alongside, same era.
H. Kamper, M. Elsner, A. Jansen, and S. Goldwater, “Unsupervised neural network based feature extraction using weak top-down constraints,” in
2015
Cited alongside, same era.
Y. Miao, H. Zhang, and F. Metze, “Speaker adaptive training of deep neural network acoustic models using i-vectors,”
2015
Cited alongside, same era.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in
2015
Cited alongside, same era.
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, and S. Khudanpur, “Deep neural network-based speaker embeddings for end-to-end speaker verification,” in
2016
Cited alongside, same era.
Y.-A. Chung, C.-C. Wu, C.-H. Shen, H.-Y. Lee, and L.-S. Lee, “Audio word2vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder,” in
2016
Cited alongside, same era.
Later among the works it cites.
2017
Later among the works it cites.
A. Jati and P. Georgiou, “Speaker2vec: Unsupervised learning and adaptation of a speaker manifold using deep neural networks with an evaluation on speaker segmentation,”
2017
Later among the works it cites.
A. Gupta, Y. Miao, L. Neves, and F. Metze, “Visual features for context-aware speech recognition,” in
2017
Later among the works it cites.
H. Bredin, “pyannote.metrics: a toolkit for reproducible evaluation, diagnostic, and error analysis of speaker diarization systems,” in
2017
Later among the works it cites.
L. McInnes, J. Healy, and S. Astels, “hdbscan: Hierarchical density based clustering,”
2017
Later among the works it cites.
2018
Closest in time.
Mozilla,
2018
Closest in time.