Fetching the paper…
Reading the bibliography…
Learning meaningful and general representations from unannotated speech that are applicable to a wide range of tasks remains challenging.
“The design for the wall street journal-based CSR corpus,”
Douglas Paul and Janet Baker, · 1992
Earlier work this paper cites.
“BLEU: A method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“Recurrent neural network based language model,”
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur, · 2010
Earlier work this paper cites.
“On the properties of neural machine translation: Encoder-decoder approaches,”
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Gaussian error linear units (gelus),”
Dan Hendrycks and Kevin Gimpel, · 2016
Earlier work this paper cites.
“Unsupervised learning of disentangled and interpretable representations from sequential data,”
Wei-Ning Hsu, Yu Zhang, and James Glass, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, et al., · 2017
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Unspeech: Unsupervised speech context embeddings,”
Benjamin Milde and Chris Biemann, · 2018
Earlier work this paper cites.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Yu-An Chung and James Glass, · 2018
Cited alongside, same era.
“Phonetic-and-semantic embedding of spoken words with applications in spoken content retrieval,”
Yi-Chen Chen, Sung-Feng Huang, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee, · 2018
Cited alongside, same era.
“Generating wikipedia by summarizing long sequences,”
Peter Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, et al., · 2018
Cited alongside, same era.
“Improving language understanding by generative pre-training,”
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever, · 2018
Cited alongside, same era.
“Deep contextualized word representations,”
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Cited alongside, same era.
“Wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Closest in time.
“Truly unsupervised acoustic word embeddings using weak top-down constraints in encoder-decoder models,”
Herman Kamper, · 2019
Closest in time.
“BERT: Pre-training of deep bidirectional Transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Closest in time.
“RoBERTa: A robustly optimized BERT pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, et al., · 2019
Closest in time.
“Towards unsupervised speech-to-text translation,”
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James Glass, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ali Kocabiyikoglu, Laurent Besacier, and Olivier Kraif, · 2018
Cited alongside, same era.
“End-to-end automatic speech translation of audiobooks,”
Alexandre Bérard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin, · 2018
Cited alongside, same era.
“Universal language model fine-tuning for text classification,”
Jeremy Howard and Sebastian Ruder, · 2018
Cited alongside, same era.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, et al., · 2018
Cited alongside, same era.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Cited alongside, same era.
“Unsupervised speech representation learning using wavenet autoencoders,”
Jan Chorowski, Ron Weiss, Samy Bengio, and Aäron van den Oord, · 2019
Cited alongside, same era.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio, · 2019
Cited alongside, same era.
Mattia Di Gangi, Matteo Negri, and Marco Turchi, · 2019
Closest in time.
“Transformers with convolutional context for ASR,”
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer, · 2019
Closest in time.
“Language modeling with deep Transformers,”
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Closest in time.
“Semi-supervised training for improving data efficiency in end-to-end speech synthesis,”
Yu-An Chung, Yuxuan Wang, Wei-Ning Hsu, Yu Zhang, and RJ Skerry-Ryan, · 2019
Closest in time.
“End-to-end text-to-speech for low-resource languages by cross-lingual transfer learning,”
Yuan-Jui Chen, Tao Tu, Cheng-Chieh Yeh, and Hung-Yi Lee, · 2019
Closest in time.
“Pre-trained text embeddings for enhanced text-to-speech synthesis,”
Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Shubham Toshniwal, and Karen Livescu, · 2019
Closest in time.
“Towards transfer learning for end-to-end speech synthesis from deep pre-trained language models,”
Wei Fang, Yu-An Chung, and James Glass, · 2019
Closest in time.