Fetching the paper…
Reading the bibliography…
We present Mockingjay as a new speech representation learning approach, where bidirectional Transformer encoders are pre-trained on a large amount of unlabeled speech.
““cloze procedure”: A new tool for measuring readability,”
Wilson L Taylor, · 1953
Earlier work this paper cites.
“Recurrent neural network based language model,”
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur, · 2010
Earlier work this paper cites.
“Adam: A method for stochastic optimization,” 2014
Diederik P. Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Audio word2vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder,” 2016
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee, · 2016
Earlier work this paper cites.
“Layer normalization,” 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton, · 2016
Earlier work this paper cites.
‘‘Attention is all you need,’’ 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Convolutional sequence to sequence learning,” 2017
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin, · 2017
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi.,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Cited alongside, same era.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Yu-An Chung and James Glass, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,” 2018
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,” 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
“Self-attentional acoustic models,”
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, and Alex Waibel, · 2018
Cited alongside, same era.
“Deep contextualized word representations,” 2018
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Closest in time.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Closest in time.
“Unsupervised end-to-end learning of discrete linguistic units for voice conversion,” 2019
Andy T. Liu, Po chun Hsu, and Hung yi Lee, · 2019
Closest in time.
“Very deep self-attention networks for end-to-end speech recognition,”
Ngoc-Quan Pham, Thai-Son Nguyen, Jan Niehues, Markus Muller, and Alex Waibel, · 2019
Closest in time.
“Roberta: A robustly optimized bert pretraining approach,” 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Cited alongside, same era.
“Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph,”
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, · 2018
Cited alongside, same era.
“Unsupervised speech representation learning using wavenet autoencoders,”
Jan Chorowski, Ron J. Weiss, Samy Bengio, and Aaron van den Oord, · 2019
Cited alongside, same era.
“Albert: A lite bert for self-supervised learning of language representations,” 2019
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut, · 2019
Closest in time.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Anonymous authors, · 2020
Closest in time.
“Unsupervised learning of efficient and robust speech representations,”
Anonymous authors, · 2020
Closest in time.