Fetching the paper…
Reading the bibliography…
We propose spoken sentence embeddings which capture both acoustic and linguistic content.
“Timit acoustic phonetic continuous speech corpus,”
John S Garofolo, · 1993
Earlier work this paper cites.
“Pitch variations and emotions in speech,”
Sylvie Mozziconacci, · 1995
Earlier work this paper cites.
“Multitask learning,”
Rich Caruana, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Efficient sampling and feature selection in whole sentence maximum entropy language models,”
Stanley F Chen and Ronald Rosenfeld, · 1999
Earlier work this paper cites.
“Visualizing data using t-sne,”
Laurens van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality,”
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean, · 2013
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“On the difficulty of training recurrent neural networks,”
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio, · 2013
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“A clockwork rnn,”
Jan Koutnik, Klaus Greff, Faustino Gomez, and Juergen Schmidhuber, · 2014
Earlier work this paper cites.
“Distributed representations of sentences and documents,”
Quoc Le and Tomas Mikolov, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Cited alongside, same era.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, · 2015
Cited alongside, same era.
“Deep unordered composition rivals syntactic methods for text classification,”
Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daumé III, · 2015
Cited alongside, same era.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Skip-thought vectors,”
Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler, · 2015
Cited alongside, same era.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
“An overview of multi-task learning in deep neural networks,”
Sebastian Ruder, · 2017
Later among the works it cites.
“Montreal forced aligner: trainable text-speech alignment using kaldi,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Later among the works it cites.
“Constituency parsing with a self-attentive encoder,”
Nikita Kitaev and Dan Klein, · 2018
Later among the works it cites.
“Learning with structured representations for negation scope extraction,”
Hao Li and Wei Lu, · 2018
Later among the works it cites.
“Deep contextualized word representations,”
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Later among the works it cites.
“Universal language model fine-tuning for text classification,”
Jeremy Howard and Sebastian Ruder, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
Dario Amodei et al., · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Cited alongside, same era.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Cited alongside, same era.
“Towards universal paraphrastic sentence embeddings,”
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu, · 2016
Cited alongside, same era.
“Char2wav: End-to-end speech synthesis,” 2017
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio, · 2017
Cited alongside, same era.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Cited alongside, same era.
Later among the works it cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Later among the works it cites.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Yu-An Chung and James Glass, · 2018
Later among the works it cites.
“An empirical evaluation of generic convolutional and recurrent networks for sequence modeling,”
Shaojie Bai, J Zico Kolter, and Vladlen Koltun, · 2018
Later among the works it cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen et al., · 2018
Later among the works it cites.
“Universal sentence encoder,”
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al., · 2018
Later among the works it cites.
“The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R Livingstone and Frank A Russo, · 2018
Later among the works it cites.
“Whole sentence neural language models,”
Yinghui Huang, Abhinav Sethy, Kartik Audhkhasi, and Bhuvana Ramabhadran, · 2018
Later among the works it cites.