Fetching the paper…
Reading the bibliography…
Query-by-example search often uses dynamic time warping (DTW) for comparing queries and proposed matching segments.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in
1992
Earlier work this paper cites.
J. Bromley, J. W. Bentz, L. Bottou, I. Guyon, Y. LeCun, C. Moore, E. Säckinger, and R. Shah, “Signature verification using a “Siamese” time delay neural network,”
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
P. Indyk and R. Motwani, “Approximate nearest neighbors: Towards removing the curse of dimensionality,” in
1998
Earlier work this paper cites.
M. Charikar, “Similarity estimation techniques from rounding algorithms,” in
2002
Earlier work this paper cites.
M. Belkin and P. Niyogi, “Laplacian eigenmaps for dimensionality reduction and data representation,”
2003
Earlier work this paper cites.
C. Allauzen, M. Mohri, and M. Saraclar, “General indexation of weighted automata: application to spoken utterance retrieval,” in
2004
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in
2005
Earlier work this paper cites.
D. R. H. Miller, M. Kleber, C. Kao, O. Kimball, T. Colthurst, S. A. Lowe, R. M. Schwartz, and H. Gish, “Rapid and accurate spoken term detection,” in
2007
Earlier work this paper cites.
L. van der Maaten and G. Hinton, “Visualizing data using t-SNE,”
2008
Earlier work this paper cites.
W. Shen, C. M. White, and T. J. Hazen, “A comparison of query-by-example methods for spoken term detection,” DTIC Document, MIT Lincoln Labs, Tech. Rep., 2009
2009
Earlier work this paper cites.
C. Parada, A. Sethy, and B. Ramabhadran, “Query-by-example spoken term detection for oov terms,” in
2009
Cited alongside, same era.
T. J. Hazen, W. Shen, and C. White, “Query-by-example spoken term detection using phonetic posteriorgram templates,” in
2009
Cited alongside, same era.
Y. Zhang and J. R. Glass, “Unsupervised spoken keyword spotting via segmental dtw on gaussian posteriorgrams,” in
2009
Cited alongside, same era.
——, “A piecewise aggregate approximation lower-bound estimate for posteriorgram-based dynamic time warping.” in
2011
Cited alongside, same era.
M. A. Carlin, S. Thomas, A. Jansen, and H. Hermansky, “Rapid evaluation of speech representations for spoken term discovery,” in
2011
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2014
Later among the works it cites.
I. Szöke, L. J. Rodríguez-Fuentes, A. Buzo, X. Anguera, F. Metze, J. Proenca, M. Lojka, and X. Xiong, “Query by example search on speech at mediaEval 2015.” in
2015
Later among the works it cites.
K. Levin, A. Jansen, and B. Van Durme, “Segmental acoustic indexing for zero resource keyword search,” in
2015
Later among the works it cites.
H. Kamper, M. Elsner, A. Jansen, and S. J. Goldwater, “Unsupervised neural network based feature extraction using weak top-down constraints,” in
2015
Later among the works it cites.
Y.-A. Chung, C.-C. Wu, C.-H. Shen, and H.-Y. Lee, “Unsupervised learning of audio segment representations using sequence-to-sequence recurrent neural networks,” in
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Jansen and B. Van Durme, “Indexing raw acoustic features for scalable zero resource search,” in
2012
Cited alongside, same era.
G. Mantena and X. Anguera, “Speed improvements to information retrieval-based dynamic time warping using hierarchical k-means clustering,” in
2013
Cited alongside, same era.
K. Levin, K. Henry, A. Jansen, and K. Livescu, “Fixed-dimensional acoustic embeddings of variable-length segments in low-resource settings,” in
2013
Cited alongside, same era.
A. Jansen, S. Thomas, and H. Hermansky, “Weak top-down constraints for unsupervised acoustic model training,” in
2013
Cited alongside, same era.
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng, “Grounded compositional semantics for finding and describing images with sentences,”
2014
Cited alongside, same era.
H. Kamper, W. Wang, and K. Livescu, “Deep convolutional acoustic word embeddings using word-pair side information,” in
2016
Later among the works it cites.
S. Settle and K. Livescu, “Discriminative acoustic word embeddings: Recurrent neural network-based approaches,” in
2016
Later among the works it cites.
H. Kamper, A. Jansen, and S. J. Goldwater, “Unsupervised word segmentation and lexicon discovery using acoustic word embeddings,”
2016
Later among the works it cites.
W. He, W. Wang, and K. Livescu, “Multi-view recurrent neural acoustic word embeddings,” in
2017
Closest in time.
K. Audhkhasi, A. Rosenberg, A. Sethy, B. Ramabhadran, and B. Kingsbury, “End-to-end ASR-free keyword search from speech,” in
2017
Closest in time.