Fetching the paper…
Reading the bibliography…
Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments.
L. R. Rabiner, A. E. Rosenberg, and S. E. Levinson, “Considerations in dynamic time warping algorithms for discrete word recognition,” IEEE Trans. Acoust., Speech, Signal Process. , vol. 26, no. 6, pp. 575–582, 1978
1978
Earlier work this paper cites.
J. Bromley, J. W. Bentz, L. Bottou, I. Guyon, Y. LeCun, C. Moore, E. Säckinger, and R. Shah, “Signature verification using a “Siamese” time delay neural network,” Int. J. Pattern Rec. , vol. 7, no. 4, pp. 669–688, 1993
1993
Earlier work this paper cites.
T. Schultz and A. Waibel, “Language-independent and language-adaptive acoustic modeling for speech recognition,” Speech Communication , vol. 35, pp. 31–51, 2001
2001
Earlier work this paper cites.
T. R. Niesler, “Language-dependent state clustering for multilingual acoustic modelling,” Speech Commun. , vol. 49, no. 6, pp. 453–463, 2007
2007
Earlier work this paper cites.
A. S. Park and J. R. Glass, “Unsupervised pattern discovery in speech,” IEEE Trans. Audio, Speech, Language Process. , vol. 16, no. 1, pp. 186–197, 2008
2008
Earlier work this paper cites.
T. J. Hazen, W. Shen, and C. White, “Query-by-example spoken term detection using phonetic posteriorgram templates,” in Proc. ASRU , 2009
2009
Earlier work this paper cites.
Y. Zhang and J. R. Glass, “Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgrams,” in Proc. ASRU , 2009
2009
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Trans. Knowl. Data Eng. , vol. 22, no. 10, pp. 1345–1359, 2009
2009
Earlier work this paper cites.
K. Q. Weinberger and L. K. Saul, “Distance metric learning for large margin nearest neighbor classification,” J. Mach. Learn. Res. , vol. 10, no. Feb, pp. 207–244, 2009
2009
Earlier work this paper cites.
G. Chechik, V. Sharma, U. Shalit, and S. Bengio, “Large scale online learning of image similarity through ranking,” J. Mach. Learn. Res. , vol. 11, pp. 1109–1135, 2010
2010
Earlier work this paper cites.
A. Jansen and B. Van Durme, “Efficient spoken term discovery using randomized algorithms,” in Proc. ASRU , 2011
2011
Earlier work this paper cites.
M. A. Carlin, S. Thomas, A. Jansen, and H. Hermansky, “Rapid evaluation of speech representations for spoken term discovery,” in Proc. Interspeech , 2011
2011
Earlier work this paper cites.
K. Veselỳ, M. Karafiát, F. Grézl, M. Janda, and E. Egorova, “The language-independent bottleneck features,” in Proc. SLT , 2012
2012
Earlier work this paper cites.
A. Jansen, E. Dupoux, S. J. Goldwater, M. Johnson, S. Khudanpur, K. Church, N. Feldman, H. Hermansky, F. Metze, R. Rose, M. Seltzer, P. Clark, I. McGraw, B. Varadarajan, E. Bennett, B. Borschinger, J. Chiu, E. Dunbar, A. Fourtassi, D. Harwath, C.-y. Lee, K. Levin, A. Norouzian, V. Peddinti, R. Richardson, T. Schatz, and S. Thomas, “A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition,” in Proc. ICASSP , 2013
2013
Earlier work this paper cites.
K. Levin, K. Henry, A. Jansen, and K. Livescu, “Fixed-dimensional acoustic embeddings of variable-length segments in low-resource settings,” in Proc. ASRU , 2013
2013
Earlier work this paper cites.
T. Schultz, N. T. Vu, and T. Schlippe, “GlobalPhone: A multilingual text & speech database in 20 languages,” in Proc. ICASSP , 2013
2013
Earlier work this paper cites.
S. Bengio and G. Heigold, “Word embeddings for speech recognition,” in Proc. Interspeech , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Levin, A. Jansen, and B. Van Durme, “Segmental acoustic indexing for zero resource keyword search,” in Proc. ICASSP , 2015
2015
Earlier work this paper cites.
C.-y. Lee, T. O’Donnell, and J. R. Glass, “Unsupervised lexicon discovery from acoustic input,” Trans. ACL , vol. 3, pp. 389–403, 2015
2015
Earlier work this paper cites.
O. J. Räsänen, G. Doyle, and M. C. Frank, “Unsupervised word discovery from speech using automatic segmentation into syllable-like units,” in Proc. Interspeech , 2015
2015
Cited alongside, same era.
F. Schroff, D. Kalenichenko, and J. Philbin, “FaceNet: a unified embedding for face recognition and clustering,” in Proc. CVPR , 2015
2015
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR , 2015
2015
Cited alongside, same era.
M. Versteegh, X. Anguera, A. Jansen, and E. Dupoux, “The Zero Resource Speech Challenge 2015: Proposed approaches and results,” in Proc. SLTU , 2016
2016
Cited alongside, same era.
Y.-A. Chung, C.-C. Wu, C.-H. Shen, and H.-Y. Lee, “Unsupervised learning of audio segment representations using sequence-to-sequence recurrent neural networks,” in Proc. Interspeech , 2016
2016
Y.-A. Chung and J. R. Glass, “Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,” in Proc. Interspeech , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Toshniwal, T. N. Sainath, R. J. Weiss, B. Li, P. Moreno, E. Weinstein, and K. Rao, “Multilingual speech recognition with a single end-to-end model,” in Proc. ICASSP , 2018
2018
Later among the works it cites.
S. Tong, P. N. Garner, and H. Bourlard, “Multilingual training and cross-lingual adaptation on ctc-based acoustic model,” Speech Commun. , vol. 104, pp. 39–46, 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
H. Kamper, W. Wang, and K. Livescu, “Deep convolutional acoustic word embeddings using word-pair side information,” in Proc. ICASSP , 2016
2016
Cited alongside, same era.
S. Settle and K. Livescu, “Discriminative acoustic word embeddings: Recurrent neural network-based approaches,” in Proc. SLT , 2016
2016
Cited alongside, same era.
S. Ghannay, Y. Estève, N. Camelin, and P. Deléglise, “Evaluation of acoustic word embeddings,” in Proc. ACL Workshop Evaluating Vector-Space Representations NLP , 2016, pp. 62–66
2016
Cited alongside, same era.
E. Dunbar, X. N. Cao, J. Benjumea, J. Karadayi, M. Bernard, L. Besacier, X. Anguera, and E. Dupoux, “The Zero Resource Speech Challenge 2017,” in Proc. ASRU , 2017
2017
Cited alongside, same era.
M. Elsner and C. Shain, “Speech segmentation with a neural encoder model of working memory,” in Proc. EMNLP , 2017
2017
Cited alongside, same era.
H. Kamper, K. Livescu, and S. Goldwater, “An embedded segmental k-means model for unsupervised segmentation and clustering of speech,” in Proc. ASRU , 2017
2017
Cited alongside, same era.
S. Settle, K. Levin, H. Kamper, and K. Livescu, “Query-by-example search with discriminative neural acoustic word embeddings,” in Proc. Interspeech , 2017
2017
Cited alongside, same era.
J. Cho, M. K. Baskar, R. Li, M. Wiesner, S. H. Mallidi, N. Yalta, M. Karafiat, S. Watanabe, and T. Hori, “Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling,” in Proc. SLT , 2018
2018
Later among the works it cites.
H. Kamper, “Truly unsupervised acoustic word embeddings using weak top-down constraints in encoder-decoder models,” in Proc. ICASSP , 2019
2019
Later among the works it cites.
H. Kamper, A. Anastassiou, and K. Livescu, “Semantic query-by-example speech search using visual grounding,” in Proc. ICASSP , 2019
2019
Later among the works it cites.
A. Haque, M. Guo, P. Verma, and L. Fei-Fei, “Audio-linguistic embeddings for spoken sentences,” in Proc. ICASSP , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Palaskar, V. Raunak, and F. Metze, “Learned in speech recognition: Contextual acoustic word embeddings,” in Proc. ICASSP , 2019
2019
Later among the works it cites.
S. Settle, K. Audhkhasi, K. Livescu, and M. Picheny, “Acoustically grounded word embeddings for improved acoustics-to-word speech recognition,” in Proc. ICASSP . IEEE, 2019, pp. 5641–5645
2019
Later among the works it cites.
M. Jung, H. Lim, J. Goo, Y. Jung, and H. Kim, “Additional shared decoder on Siamese multi-view encoders for learning acoustic word embeddings,” in Proc. ASRU , 2019
2019
Later among the works it cites.
R. Menon, H. Kamper, E. Van Der Westhuizen, J. Quinn, and T. R. Niesler, “Feature exploration for almost zero-resource asr-free keyword spotting using a multilingual bottleneck extractor and correspondence autoencoders,” in Proc. Interspeech , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
O. Adams, M. Wiesner, S. Watanabe, and D. Yarowsky, “Massively multilingual adversarial speech recognition,” in Proc. ACL , 2019
2019
Later among the works it cites.
S. Ruder, “Neural transfer learning for natural language processing,” Ph.D. dissertation, NUI Galway, Ireland, 2019
2019
Later among the works it cites.
H. Kamper, Y. Matusevych, and S. J. Goldwater, “Multilingual acoustic word embedding models for processing zero-resource languages,” in Proc. ICASSP , 2020
2020
Closest in time.
Y. Matusevych, H. Kamper, and S. Goldwater, “Analyzing autoencoder-based acoustic word embeddings,” in BAICS Workshop ICLR , 2020
2020
Closest in time.
E. Hermann, H. Kamper, and S. J. Goldwater, “Multilingual and unsupervised subword modeling for zero resource languages,” Comput. Speech Language , 2020
2020
Closest in time.