Fetching the paper…
Reading the bibliography…
Unsupervised discovery of acoustic tokens from audio corpora without annotation and learning vector representations for these tokens have been widely studied.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,”
1993
Earlier work this paper cites.
A. S. Park and J. R. Glass, “Unsupervised pattern discovery in speech,”
2008
Earlier work this paper cites.
N. Dehak, R. Dehak, P. Kenny, N. Brümmer, P. Ouellet, and P. Dumouchel, “Support vector machines versus fast scoring in the low-dimensional total variability space for speaker verification,” in
2009
Earlier work this paper cites.
B. Schuller, S. Steidl, and A. Batliner, “The interspeech 2009 emotion challenge,” in
2009
Earlier work this paper cites.
J. Glass, “Towards unsupervised speech processing,” in
2012
Earlier work this paper cites.
J. Driesen and H. Van hamme, “Fast word acquisition in an NMF-based learning framework,” in
2012
Earlier work this paper cites.
H. Wang, C.-C. Leung, T. Lee, B. Ma, and H. Li, “An acoustic segment modeling approach to query-by-example spoken term detection,” in
2012
Earlier work this paper cites.
Y. Zhang, R. Salakhutdinov, H.-A. Chang, and J. Glass, “Resource configurable spoken query detection using deep boltzmann machines,” in
2012
Earlier work this paper cites.
A. Norouzian, A. Jansen, R. C. Rose, and S. Thomas, “Exploiting discriminative point process models for spoken term detection,” in
2012
Earlier work this paper cites.
K. Levin, K. Henry, A. Jansen, and K. Livescu, “Fixed-dimensional acoustic embeddings of variable-length segments in low-resource settings,” in
2013
Earlier work this paper cites.
H.-y. Lee and L.-s. Lee, “Enhanced spoken term detection using support vector machines and weighted pseudo examples,”
2013
Earlier work this paper cites.
I.-F. Chen and C.-H. Lee, “A hybrid hmm/dnn approach to keyword spotting of short words.” in
2013
Earlier work this paper cites.
K. Cho, B. van Merriënboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” in
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in
2014
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Cited alongside, same era.
C. J. Maddison, D. Tarlow, and T. Minka, “A* sampling,” in
2014
Cited alongside, same era.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in
2015
Cited alongside, same era.
2017
Later among the works it cites.
H. Kamper, A. Jansen, and S. Goldwater, “A segmental framework for fully-unsupervised large-vocabulary speech recognition,”
2017
Later among the works it cites.
C.-T. Chung, C.-Y. Tsai, C.-H. Liu, and L.-S. Lee, “Unsupervised iterative deep learning of speech features and acoustic tokens with applications to spoken term detection,”
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Lyzinski, G. Sell, and A. Jansen, “An evaluation of graph clustering methods for unsupervised term discovery,” in
2015
Cited alongside, same era.
K. Levin, A. Jansen, and B. Van Durme, “Segmental acoustic indexing for zero resource keyword search,” in
2015
Cited alongside, same era.
H. Kamper, W. Wang, and K. Livescu, “Deep convolutional acoustic word embeddings using word-pair side information,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,”
2016
Cited alongside, same era.
2016
Cited alongside, same era.
O. Bojar, R. Chatterjee, C. Federmann, Y. Graham, B. Haddow, M. Huck, A. J. Yepes, P. Koehn, V. Logacheva, C. Monz
2016
Cited alongside, same era.
2017
Later among the works it cites.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein gan,”
2017
Later among the works it cites.
L. Yu, W. Zhang, J. Wang, and Y. Yu, “Seqgan: Sequence generative adversarial nets with policy gradient.” 2017
2017
Later among the works it cites.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of wasserstein gans,” in
2017
Later among the works it cites.
C.-T. Chung and L.-S. Lee, “Unsupervised discovery of structured acoustic tokens with applications to spoken term detection,”
2018
Closest in time.
A. Conneau, G. Lample, M. Ranzato, L. Denoyer, and H. Jégou, “Word translation without parallel data,” in
2018
Closest in time.
G. Lample, L. Denoyer, and M. Ranzato, “Unsupervised machine translation using monolingual corpora only,” in
2018
Closest in time.
Y.-H. Wang, H.-y. Lee, and L.-s. Lee, “Segmental audio word2vec: Representing utterances as sequences of vectors with applications in spoken term detection,” in
2018
Closest in time.