Fetching the paper…
Reading the bibliography…
Previous researches on acoustic word embeddings used in query-by-example spoken term detection have shown remarkable performance improvements when using a triplet network.
“Dynamic programming algorithm optimization for spoken word recognition,”
H. Sakoe and S. Chiba, · 1978
Earlier work this paper cites.
“Considerations in dynamic time warping algorithms for discrete word recognition,”
L. R. Rabiner, A. Rosenberg, and S. Levinson, · 1978
Earlier work this paper cites.
“The DARPA 1000-word resource management database for continuous speech recognition,”
P. Price, W. M. Fisher, J. Bernstein, and D. S. Pallett, · 1988
Earlier work this paper cites.
“A hidden markov model based keyword recognition system,”
R. C. Rose and D. B. Paul, · 1990
Earlier work this paper cites.
“Automatic recognition of keywords in unconstrained speech using hidden markov models,”
J. G. Wilpon, L. R. Rabiner, C. H. Lee, and E. R. Goldman, · 1990
Earlier work this paper cites.
“Improvements and applications for key word recognition using hidden markov modeling techniques,”
J. G. Wilpon, L. G. Miller, and P. Modi, · 1991
Earlier work this paper cites.
“The design for the wall street journal-based CSR corpus,”
D. B. Paul and J. M. Baker, · 1992
Earlier work this paper cites.
“Signature verification using a “siamese” time delay neural network,”
J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah, · 1994
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
M. Schuster and K. K. Paliwal, · 1997
Earlier work this paper cites.
“Visualizing data using t-SNE,”
L. Van Der Maaten and G. Hinton, · 2008
Earlier work this paper cites.
“Unsupervised spoken keyword spotting via segmental DTW on gaussian posteriorgrams,”
Y. Zhang and J. R. Glass, · 2009
Earlier work this paper cites.
“Query-by-example spoken term detection using phonetic posteriorgram templates,”
T. J. Hazen, W. Shen, and C. White, · 2009
Earlier work this paper cites.
“ “your word is my command”: Google search by voice: A case study,”
J. Schalkwyk · 2010
Cited alongside, same era.
“Rapid evaluation of speech representations for spoken term discovery,”
M. A. Carlin, S. Thomas, A. Jansen, and H. Hermansky, · 2011
Cited alongside, same era.
“The Kaldi speech recognition toolkit,”
D. Povey · 2011
Cited alongside, same era.
“Fixed-dimensional acoustic embeddings of variable-length segments in low-resource settings,”
K. Levin, K. Henry, A. Jansen, and K. Livescu, · 2013
Cited alongside, same era.
“Small-footprint keyword spotting using deep neural networks,”
G. Chen, C. Parada, and G. Heigold, · 2014
Cited alongside, same era.
“Query-by-example keyword spotting using long short-term memory networks,”
G. Chen, C. Parada, and T. N. Sainath, · 2015
Cited alongside, same era.
“Investigating neural network based query-by-example keyword spotting approach for personalized wake-up word detection in mandarin chinese,”
J. Hou, L. Xie, and Z. Fu, · 2016
Later among the works it cites.
“Deep convolutional acoustic word embeddings using word-pair side information,”
H. Kamper, W. Wang, and K. Livescu, · 2016
Later among the works it cites.
“Discriminative acoustic word embeddings: Recurrent neural network-based approaches,”
S. Settle and K. Livescu, · 2016
Later among the works it cites.
“CNN-based bottleneck feature for noise robust query-by-example spoken term detection,”
H. Lim, Y. Kim, Y. Kim, and H. Kim, · 2017
Later among the works it cites.
“Query-by-example search with discriminative neural acoustic word embeddings,”
S. Settle, K. Levin, H. Kamper, and K. Livescu, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deep metric learning using triplet network,”
E. Hoffer and N. Ailon, · 2015
Cited alongside, same era.
“Convolutional neural networks for small-footprint keyword spotting,”
T. N. Sainath and C. Parada, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2015
Cited alongside, same era.
“TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015,
M. Abadi · 2015
Cited alongside, same era.
“Facenet: A unified embedding for face recognition and clustering,”
F. Schroff, D. Kalenichenko, and J. Philbin, · 2015
Cited alongside, same era.
“Multi-task learning and weighted cross-entropy for DNN-based keyword spotting,”
S. Panchapagesan, M. Sun, A. Khare, S. Matsoukas, A. Mandal, B. Hoffmeister, and S. Vitaladevuni, · 2016
Cited alongside, same era.
“Multitask learning with low-level auxiliary tasks for encoder-decoder based speech recognition,”
S. Toshniwal, H. Tang, L. Lu, and K. Livescu, · 2017
Later among the works it cites.
“Simple triplet loss based on intra/inter-class metric learning for face verification,”
Z. Ming, J. Chazalon, M. M. Luqman, M. Visani, and J. C. Burie, · 2017
Later among the works it cites.
“Segmental audio word2vec: Representing utterances as sequences of vectors with applications in spoken term detection,”
Y. Wang, H. Lee, and L. Lee, · 2018
Closest in time.
“Hierarchical multitask learning for CTC-based speech recognition,”
K. Krishna, S. Toshniwal, and K. Livescu, · 2018
Closest in time.
“Learning acoustic word embeddings with temporal context for query-by-example speech search,”
Y. Yougen, L. Cheung-Chi, X. Lei, C. Hongjie, M. Bin, and L. Haizhou, · 2018
Closest in time.
“Attention-based end-to-end models for small-footprint keyword spotting,”
S. Changhao, Z. Junbo, W. Yujun, and X. Lei, · 2041
Closest in time.