Fetching the paper…
Reading the bibliography…
Word embedding or Word2Vec has been successful in offering semantics for text words learned from the context of words.
“Leveraging relevance cues for improved spoken document retrieval,”
Pei-Ning Chen, Kuan-Yu Chen, and Berlin Chen, · 2011
Earlier work this paper cites.
“Spoken document retrieval with unsupervised query modeling techniques,”
Berlin Chen, Kuan-Yu Chen, Pei-Ning Chen, and Yi-Wen Chen, · 2012
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality,”
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean, · 2013
Earlier work this paper cites.
“Fixed-dimensional acoustic embeddings of variable-length segments in low-resource settings,”
Keith Levin, Katharine Henry, Aren Jansen, and Karen Livescu, · 2013
Earlier work this paper cites.
“Enhancing query expansion for semantic retrieval of spoken content with automatically discovered acoustic patterns,”
Hung-yi Lee, Yun-Chiao Li, Cheng-Tao Chung, and Lin-shan Lee, · 2013
Earlier work this paper cites.
“Interactive spoken content retrieval by extended query model and continuous state space markov decision process,”
Tsung-Hsien Wen, Hung-Yi Lee, Pei-hao Su, and Lin-Shan Lee, · 2013
Earlier work this paper cites.
“Glove: Global vectors for word representation,”
Jeffrey Pennington, Richard Socher, and Christopher Manning, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Earlier work this paper cites.
“Word embeddings for speech recognition,”
Samy Bengio and Georg Heigold, · 2014
Earlier work this paper cites.
“Learning phrase representations using rnn encoder-decoder for statistical machine translation,”
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Improved transition-based parsing by modeling characters instead of words with lstms,”
Miguel Ballesteros, Chris Dyer, and Noah A Smith, · 2015
Earlier work this paper cites.
“Effective approaches to attention-based neural machine translation,”
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning, · 2015
Earlier work this paper cites.
“Sequence to sequence - video to text,”
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond J. Mooney, Trevor Darrell, and Kate Saenko, · 2015
Earlier work this paper cites.
“Deep convolutional inverse graphics network,”
Tejas D. Kulkarni, Will Whitney, Pushmeet Kohli, and Joshua B. Tenenbaum, · 2015
Earlier work this paper cites.
“Spoken content retrieval—beyond cascading speech recognition with text retrieval,”
Lin-shan Lee, James Glass, Hung-yi Lee, and Chun-an Chan, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Enriching word vectors with subword information,”
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov, · 2016
Cited alongside, same era.
“Neural architectures for named entity recognition,”
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer, · 2016
Cited alongside, same era.
Barbara Plank, Anders Søgaard, and Yoav Goldberg, · 2016
Cited alongside, same era.
“Character-aware neural language models.,”
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush, · 2016
“Query-by-example search with discriminative neural acoustic word embeddings,”
Shane Settle, Keith Levin, Herman Kamper, and Karen Livescu, · 2017
Later among the works it cites.
“Parsing speech: A neural approach to integrating lexical and acoustic-prosodic information,”
Trang Tran, Shubham Toshniwal, Mohit Bansal, Kevin Gimpel, Karen Livescu, and Mari Ostendorf, · 2017
Later among the works it cites.
“End-to-end neural segmental models for speech recognition,”
Hao Tang, Liang Lu, Lingpeng Kong, Kevin Gimpel, Karen Livescu, Chris Dyer, Noah A Smith, and Steve Renals, · 2017
Later among the works it cites.
“An embedded segmental k-means model for unsupervised segmentation and clustering of speech,”
Herman Kamper, Karen Livescu, and Sharon Goldwater, · 2017
Later among the works it cites.
“A segmental framework for fully-unsupervised large-vocabulary speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Multi-view recurrent neural acoustic word embeddings,”
Wanjia He, Weiran Wang, and Karen Livescu, · 2016
Cited alongside, same era.
“Discriminative acoustic word embeddings: Recurrent neural network-based approaches,”
Shane Settle and Karen Livescu, · 2016
Cited alongside, same era.
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee, · 2016
Cited alongside, same era.
“Deep convolutional acoustic word embeddings using word-pair side information,”
Herman Kamper, Weiran Wang, and Karen Livescu, · 2016
Cited alongside, same era.
“Infogan: Interpretable representation learning by information maximizing generative adversarial nets,”
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel, · 2016
Cited alongside, same era.
“beta-vae: Learning basic visual concepts with a constrained variational framework,”
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner, · 2016
Cited alongside, same era.
Konstantinos Bousmalis, George Trigeorgis, Nathan Silberman, Dilip Krishnan, and Dumitru Erhan, · 2016
Cited alongside, same era.
Herman Kamper, Aren Jansen, and Sharon Goldwater, · 2017
Later among the works it cites.
“Unsupervised learning of disentangled and interpretable representations from sequential data,”
Wei-Ning Hsu, Yu Zhang, and James R. Glass, · 2017
Later among the works it cites.
“Unsupervised adaptation with domain separation networks for robust speech recognition,”
Zhong Meng, Zhuo Chen, Vadim Mazalov, Jinyu Li, and Yifan Gong, · 2017
Later among the works it cites.
“Unsupervised learning of semantic audio representations,”
Aren Jansen, Manoj Plakal, Ratheet Pandya, Daniel PW Ellis, Shawn Hershey, Jiayang Liu, R Channing Moore, and Rif A Saurous, · 2017
Later among the works it cites.
“Towards learning semantic audio representations from unlabeled data,”
Aren Jansen, Manoj Plakal, Ratheet Pandya, Daniel PW Ellis, Shawn Hershey, Jiayang Liu, R Channing Moore, and Rif A Saurous, · 2017
Later among the works it cites.
“Word translation without parallel data,”
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou, · 2017
Later among the works it cites.
“Improved training of wasserstein gans,”
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville, · 2017
Later among the works it cites.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Yu-An Chung and James R. Glass, · 2018
Closest in time.
“Towards unsupervised automatic speech recognition trained by unaligned speech and text only,”
Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang, and Hung-yi Lee, · 2018
Closest in time.
“Segmental audio word2vec: Representing utterances as sequences of vectors with applications in spoken term detection,”
Yu-Hsuan Wang, Hung-yi Lee, and Lin-shan Lee, · 2018
Closest in time.
“An iterative closest point method for unsupervised word translation,”
Yedid Hoshen and Lior Wolf, · 2018
Closest in time.