Fetching the paper…
Reading the bibliography…
Self-supervised speech representations have been shown to be effective in a variety of speech applications.
“The design for the Wall Street Journal-based CSR corpus,”
Douglas Paul and Janet Baker, · 1992
Earlier work this paper cites.
“Efficient estimation of word representations in vector space,”
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean, · 2013
Earlier work this paper cites.
“LibriSpeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron van den Oord, Oriol Vinyals, et al., · 2017
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Unspeech: Unsupervised speech context embeddings,”
Benjamin Milde and Chris Biemann, · 2018
Earlier work this paper cites.
“Deep contextualized word representations,”
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Earlier work this paper cites.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Cited alongside, same era.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional Transformersfor language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Unsupervised speech representation learning using wavenet autoencoders,”
Jan Chorowski, Ron Weiss, Samy Bengio, and Aäron van den Oord, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Closest in time.
“Unsupervised pre-training of bidirectional speech encoders via masked reconstruction,”
Weiran Wang, Qingming Tang, and Karen Livescu, · 2020
Closest in time.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional Transformer encoders,”
Andy Liu, Shu-Wen Yang, Po-Han Chi, Po-Chun Hsu, and Hung-Yi Lee, · 2020
Closest in time.
“Improved speech representations with multi-target autoregressive predictive coding,”
Yu-An Chung and James Glass, · 2020
Closest in time.
“Vector-quantized autoregressive predictive coding,”
Yu-An Chung, Hao Tang, and James Glass, · 2020
Closest in time.
“Deep contextualized acoustic representations for semi-supervised speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, et al., · 2019
Cited alongside, same era.
“Unsupervised pretraining transfers well across languages,”
Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, and Emmanuel Dupoux, · 2020
Cited alongside, same era.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2020
Cited alongside, same era.
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff, · 2020
Closest in time.
“SpeechBERT: Cross-modal pre-trained language model for end-to-end spoken question answering,”
Yung-Sung Chuang, Chi-Liang Liu, and Hung-Yi Lee, · 2020
Closest in time.
“Speech-XLNet: Unsupervised acoustic model pretraining for self-attention networks,”
Xingchen Song, Guangsen Wang, Zhiyong Wu, Yiheng Huang, Dan Su, Dong Yu, and Helen Meng, · 2020
Closest in time.