Fetching the paper…
Reading the bibliography…
Acoustic word embeddings --- fixed-dimensional vector representations of variable-length spoken word segments --- have begun to be considered for tasks such as speech recognition and query-by-example search.
“Switchboard: Telephone speech corpus for research and development,”
John J Godfrey, Edward C Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“Signature verification using a ‘Siamese’ time delay neural network,”
Jane Bromley, James W Bentz, Léon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard Säckinger, and Roopak Shah, · 1993
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Moving Beyond the ’Beads-on-a-String’ Model of Speech,”
M. Ostendorf, · 1999
Earlier work this paper cites.
Proceedings of the International Tutorial and Research Workshop on Pronunciation Modeling and Lexicon Adaptation for Spoken Language Technology
ISCA, · 2002
Earlier work this paper cites.
“Template-based continuous speech recognition,”
Mathias De Wachter, Mike Matton, Kris Demuynck, Patrick Wambacq, Ronald Cools, and Dirk Van Compernolle, · 2007
Earlier work this paper cites.
“Visualizing data using t-sne,”
Laurens van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“A segmental CRF approach to large vocabulary continuous speech recognition,”
Geoffrey Zweig and Patrick Nguyen, · 2009
Earlier work this paper cites.
“Rectified linear units improve restricted boltzmann machines,”
Vinod Nair and Geoffrey E Hinton, · 2010
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization,”
John Duchi, Elad Hazan, and Yoram Singer, · 2010
Earlier work this paper cites.
“Rapid evaluation of speech representations for spoken term discovery,”
Michael A Carlin, Samuel Thomas, Aren Jansen, and Hynek Hermansky, · 2011
Earlier work this paper cites.
“Torch7: A matlab-like environment for machine learning,”
Ronan Collobert, Koray Kavukcuoglu, and Clément Farabet, · 2011
Earlier work this paper cites.
“Subword modeling for automatic speech recognition: Past, present, and emerging approaches,”
Karen Livescu, Eric Fosler-Lussier, and Florian Metze, · 2012
Earlier work this paper cites.
“Investigations on exemplar-based features for speech recognition towards thousands of hours of unsupervised, noisy data,”
Georg Heigold, Patrick Nguyen, Mitchel Weintraub, and Vincent Vanhoucke, · 2012
Earlier work this paper cites.
“Speaker independent discriminant feature extraction for acoustic pattern-matching,”
Xavier Anguera, · 2012
Cited alongside, same era.
“Fast spoken query detection using lower-bound dynamic time warping on graphical processing units,”
Yaodong Zhang, Kiarash Adl, and James Glass, · 2012
Cited alongside, same era.
“Word-level acoustic modeling with convolutional vector regression,”
Andrew L Maas, Stephen D Miller, Tyler M O’neil, Andrew Y Ng, and Patrick Nguyen, · 2012
Cited alongside, same era.
“ADADELTA: an adaptive learning rate method,”
Matthew D. Zeiler, · 2012
Cited alongside, same era.
“The spoken web search task at MediaEval 2012,”
Florian Metze, Xavier Anguera, Etienne Barnard, Marelie Davel, and Guillaume Gravier, · 2013
Cited alongside, same era.
“Dropout: a simple way to prevent neural networks from overfitting.,”
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Later among the works it cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2014
Later among the works it cites.
“Coping with channel mismatch in query-by-example - BUT QUESST 2014,”
Igor Szöke, Miroslav Skácel, Lukás̆ Burget, and Jan “Honza” C̆ernocký, · 2015
Later among the works it cites.
“Segmental acoustic indexing for zero resource keyword search,”
Keith Levin, Aren Jansen, and Benjamin Van Durme, · 2015
Later among the works it cites.
“Query-by-example keyword spotting using long short-term memory networks,”
Guoguo Chen, Carolina Parada, and Tara N Sainath, · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keith Levin, Katharine Henry, Aren Jansen, and Karen Livescu, · 2013
Cited alongside, same era.
“Weak top-down constraints for unsupervised acoustic model training,”
Aren Jansen, Samuel Thomas, and Hynek Hermansky, · 2013
Cited alongside, same era.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Cited alongside, same era.
“Word embeddings for speech recognition,”
Samy Bengio and Georg Heigold, · 2014
Cited alongside, same era.
“Word-level invariant representations from acoustic waveforms.,”
Stephen Voinea, Chiyuan Zhang, Georgios Evangelopoulos, Lorenzo Rosasco, and Tomaso Poggio, · 2014
Cited alongside, same era.
“Unsupervised lexical clustering of speech segments using fixed-dimensional acoustic embeddings,”
Herman Kamper, Aren Jansen, Simon King, and Sharon Goldwater, · 2014
Cited alongside, same era.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Fully unsupervised small-vocabulary speech recognition using a segmental bayesian model,”
Herman Kamper, Aren Jansen, and Sharon Goldwater, · 2015
Later among the works it cites.
“Unsupervised neural network based feature extraction using weak top-down constraints,”
H. Kamper, M. Elsner, A. Jansen, and S. J. Goldwater, · 2015
Later among the works it cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Later among the works it cites.
“A study of the recurrent neural network encoder-decoder for large vocabulary speech recognition,”
Liang Lu, Xingxing Zhang, Kyunghyun Cho, and Steve Renals, · 2015
Later among the works it cites.
“rnn : Recurrent library for torch,”
Nicholas Léonard, Sagar Waghmare, Yang Wang, and Jin-Hwa Kim, · 2015
Later among the works it cites.
“Deep convolutional acoustic word embeddings using word-pair side information,”
Herman Kamper, Weiran Wang, and Karen Livescu, · 2016
Closest in time.
“Unsupervised learning of audio segment representations using sequence-to-sequence recurrent neural networks,”
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, and Hung-Yi Lee, · 2016
Closest in time.
“Learning neural network representations using cross-lingual bottleneck features with word-pair information,”
Yougen Yuan, Cheung-Chi Leung, Lei Xie, Bin Ma, and Haizhou Li, · 2016
Closest in time.