Fetching the paper…
Reading the bibliography…
Training objectives based on predictive coding have recently been shown to be very effective at learning meaningful representations from unlabeled speech.
Improving Transformer-based speech recognition using unsupervised pre-training
Dongwei Jiang, Xiaoning Lei, Wubo Li, Ne Luo, Yuxuan Hu, et al. 2019 · 1910
Earlier work this paper cites.
Speech-XLNet: Unsupervised acoustic model pretraining for self-attention networks
Xingchen Song, Guangsen Wang, Zhiyong Wu, Yiheng Huang, Dan Su, et al. 2019 · 1910
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed. 2019 · 1911
Earlier work this paper cites.
Speaker-independent phone recognition using hidden markov models
Kai-Fu Lee and Hsiao-Wuen Hon. 1989 · 1989
Earlier work this paper cites.
The design for the wall street journal-based CSR corpus
Douglas Paul and Janet Baker. 1992 · 1992
Earlier work this paper cites.
DARPA TIMIT acoustic-phonetic continuous speech corpus
John Garofolo, Lori Lamel, William Fisher, Jonathan Fiscus, David Pallett, and Nancy Dahlgren. 1993 · 1993
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Attention-based models for speech recognition
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Cited alongside, same era.
Audio word2vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc Le, Mohammad Norouzi, et al. 2016 · 2016
Cited alongside, same era.
Unsupervised learning of disentangled and interpretable representations from sequential data
Wei-Ning Hsu, Yu Zhang, and James Glass. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Unspeech: Unsupervised speech context embeddings
Benjamin Milde and Chris Biemann. 2018 · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Later among the works it cites.
An unsupervised autoregressive model for speech representation learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019 · 2019
Later among the works it cites.
Truly unsupervised acoustic word embeddings using weak top-down constraints in encoder-decoder models
Herman Kamper. 2019 · 2019
Later among the works it cites.
Learning problem-agnostic speech representations from multiple self-supervised tasks
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Phonetic-and-semantic embedding of spoken words with applications in spoken content retrieval
Yi-Chen Chen, Sung-Feng Huang, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee. 2018 · 2018
Cited alongside, same era.
Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech
Yu-An Chung and James Glass. 2018 · 2018
Cited alongside, same era.
Unsupervised cross-modal alignment of speech and text embedding spaces
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James Glass. 2018 · 2018
Cited alongside, same era.
Augmenting LibriSpeech with French translations: A multimodal corpus for direct speech translation evaluation
Ali Kocabiyikoglu, Laurent Besacier, and Olivier Kraif. 2018 · 2018
Cited alongside, same era.
Generating wikipedia by summarizing long sequences
Peter Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018 · 2018
Cited alongside, same era.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 2019
Later among the works it cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2020 · 2020
Closest in time.
Generative pre-training for speech with autoregressive predictive coding
Yu-An Chung and James Glass. 2020 · 2020
Closest in time.
Mockingjay: Unsupervised speech representation learning with deep bidirectional Transformer encoders
Andy Liu, Shu-Wen Yang, Po-Han Chi, Po-Chun Hsu, and Hung-Yi Lee. 2020 · 2020
Closest in time.
Unsupervised speech representation learning using wavenet autoencoders
Jan Chorowski, Ron Weiss, Samy Bengio, and Aäron van den Oord. 2019 · 2053
Closest in time.