Fetching the paper…
Reading the bibliography…
We investigate unsupervised models that can map a variable-duration speech segment to a fixed-dimensional representation.
“Considerations in dynamic time warping algorithms for discrete word recognition,”
L. R. Rabiner, A. E. Rosenberg, and S. E. Levinson, · 1978
Earlier work this paper cites.
“Supervised neural networks for the classification of structures,”
A. Sperduti and A. Starita, · 1997
Earlier work this paper cites.
“The Buckeye corpus of conversational speech: Labeling conventions and a test of transcriber reliability,”
M. A. Pitt, K. Johnson, E. Hume, S. Kiesling, and W. Raymond, · 2005
Earlier work this paper cites.
“Unsupervised pattern discovery in speech,”
A. S. Park and J. R. Glass, · 2008
Earlier work this paper cites.
“Visualizing data using t-SNE,”
L. Van der Maaten and G. Hinton, · 2008
Earlier work this paper cites.
“Query-by-example spoken term detection using phonetic posteriorgram templates,”
T. J. Hazen, W. Shen, and C. White, · 2009
Earlier work this paper cites.
“Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgrams,”
Y. Zhang and J. R. Glass, · 2009
Earlier work this paper cites.
“Efficient spoken term discovery using randomized algorithms,”
A. Jansen and B. Van Durme, · 2011
Earlier work this paper cites.
“Rapid evaluation of speech representations for spoken term discovery,”
M. A. Carlin, S. Thomas, A. Jansen, and H. Hermansky, · 2011
Earlier work this paper cites.
“Computational modeling of phonetic and lexical learning in early language acquisition: Existing models and future directions,”
O. J. Räsänen, · 2012
Earlier work this paper cites.
“Word-level acoustic modeling with convolutional vector regression,”
A. L. Maas, S. D. Miller, T. M. O’Neil, A. Y. Ng, and P. Nguyen, · 2012
Earlier work this paper cites.
“A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition,”
A. Jansen et al., · 2013
Earlier work this paper cites.
“Fixed-dimensional acoustic embeddings of variable-length segments in low-resource settings,”
K. Levin, K. Henry, A. Jansen, and K. Livescu, · 2013
Earlier work this paper cites.
“Auto-encoding variational bayes,”
D. P. Kingma and M. Welling, · 2013
Earlier work this paper cites.
“Learning phrase representations using RNN encoder-decoder for statistical machine translation,”
K. Cho et al., · 2014
Cited alongside, same era.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, · 2014
Cited alongside, same era.
“Learning grounded meaning representations with autoencoders,”
C. Silberer and M. Lapata, · 2014
Cited alongside, same era.
“A smartphone-based ASR data collection tool for under-resourced languages,”
N. J. De Vries et al., · 2014
Cited alongside, same era.
“Segmental acoustic indexing for zero resource keyword search,”
K. Levin, A. Jansen, and B. Van Durme, · 2015
Cited alongside, same era.
“Unsupervised lexicon discovery from acoustic input,”
“The Zero Resource Speech Challenge 2017,”
E. Dunbar et al., · 2017
Later among the works it cites.
“Speech segmentation with a neural encoder model of working memory,”
M. Elsner and C. Shain, · 2017
Later among the works it cites.
“An embedded segmental k-means model for unsupervised segmentation and clustering of speech,”
H. Kamper, K. Livescu, and S. Goldwater, · 2017
Later among the works it cites.
“Unsupervised speech signal to symbol transformation for zero resource speech applications,”
S. Bhati, S. Nayak, and K. S. R. Murty, · 2017
Later among the works it cites.
“End-to-end ASR-free keyword search from speech,”
K. Audhkhasi, A. Rosenberg, A. Sethy, B. Ramabhadran, and B. Kingsbury, · 2017
Later among the works it cites.
“Query-by-example search with discriminative neural acoustic word embeddings,”
S. Settle, K. Levin, H. Kamper, and K. Livescu, · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C.-y. Lee, T. O’Donnell, and J. R. Glass, · 2015
Cited alongside, same era.
“Unsupervised word discovery from speech using automatic segmentation into syllable-like units,”
O. J. Räsänen, G. Doyle, and M. C. Frank, · 2015
Cited alongside, same era.
“A comparison of neural network methods for unsupervised representation learning on the Zero Resource Speech Challenge,”
D. Renshaw, H. Kamper, A. Jansen, and S. J. Goldwater, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2015
Cited alongside, same era.
“Breaking the unwritten language barrier: The BULB project,”
G. Adda et al., · 2016
Cited alongside, same era.
“The Zero Resource Speech Challenge 2015: Proposed approaches and results,”
M. Versteegh, X. Anguera, A. Jansen, and E. Dupoux, · 2016
Cited alongside, same era.
“Unsupervised learning of audio segment representations using sequence-to-sequence recurrent neural networks,”
Y.-A. Chung, C.-C. Wu, C.-H. Shen, and H.-Y. Lee, · 2016
Cited alongside, same era.
Later among the works it cites.
“Unsupervised transformation learning via convex relaxations,”
T. B. Hashimoto, P. S. Liang, and J. C. Duchi, · 2017
Later among the works it cites.
“Unsupervised learning of disentangled and interpretable representations from sequential data,”
W.-N. Hsu, Y. Zhang, and J. R. Glass, · 2017
Later among the works it cites.
“Segmental audio word2vec: Representing utterances as sequences of vectors with applications in spoken term detection,”
Y.-H. Wang, H.-y. Lee, and L.-s. Lee, · 2018
Closest in time.
“Unsupervised cross-modal alignment of speech and text embedding spaces,”
Y.-A. Chung, W.-H. Weng, S. Tong, and J. R. Glass, · 2018
Closest in time.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Y.-A. Chung and J. R. Glass, · 2018
Closest in time.
“Learning word embeddings: unsupervised methods for fixed-size representations of variable-length speech segments,”
N. Holzenberger, M. Du, J. Karadayi, R. Riad, and E. Dupoux, · 2018
Closest in time.
“Phonetic-and-semantic embedding of spoken words with applications in spoken content retrieval,”
Y.-C. Chen, S.-F. Huang, C.-H. Shen, H.-y. Lee, and L.-s. Lee, · 2018
Closest in time.