Fetching the paper…
Reading the bibliography…
For self-supervised speech processing, it is crucial to use pretrained models as speech representation extractors.
“Visualizing data using t-sne,”
Maaten, Laurens van der and Hinton, Geoffrey, · 2008
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Using the output embedding to improve language models,”
Ofir Press and Lior Wolf, · 2017
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Earlier work this paper cites.
“Deep contextualized word representations,”
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Earlier work this paper cites.
“Improving language understanding by generative pre-training,”
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, · 2019
Earlier work this paper cites.
“Improving transformer-based speech recognition using unsupervised pre-training,”
Dongwei Jiang, Xiaoning Lei, Wubo Li, Ne Luo, Yuxuan Hu, Wei Zou, and Xiangang Li, · 2019
Cited alongside, same era.
“Semi-Supervised Sequence-to-Sequence ASR Using Unpaired Speech and Text,”
Murali Karthick Baskar, Shinji Watanabe, Ramon Astudillo, Takaaki Hori, Lukáš Burget, and Jan Černocký, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised Pre-Training for Speech Recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“Albert: A lite bert for self-supervised learning of language representations,”
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut, · 2019
Cited alongside, same era.
“Xlnet: Generalized autoregressive pretraining for language understanding,”
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le, · 2019
Cited alongside, same era.
“Recurrent stacking of layers for compact neural machine translation models,”
Raj Dabre and Atsushi Fujita, · 2019
Later among the works it cites.
“Universal transformers,” Nov. 21 2019,
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob D Uszkoreit, and Lukasz Mieczyslaw Kaiser, · 2019
Later among the works it cites.
“What does bert learn about the structure of language?,”
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah, · 2019
Later among the works it cites.
“Analyzing Phonetic and Graphemic Representations in End-to-End Automatic Speech Recognition,”
Yonatan Belinkov, Ahmed Ali, and James Glass, · 2019
Later among the works it cites.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee, · 2020
Closest in time.
“Deep contextualized acoustic representations for semi-supervised speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Roberta: A robustly optimized bert pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Cited alongside, same era.
“An Unsupervised Autoregressive Model for Speech Representation Learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Cited alongside, same era.
“Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks,”
Xingchen Song, Guangsen Wang, Zhiyong Wu, Yiheng Huang, Dan Su, Dong Yu, and Helen Meng, · 2019
Cited alongside, same era.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2019
Cited alongside, same era.
“Sharing attention weights for fast transformer,”
Tong Xiao, Yinqiao Li, Jingbo Zhu, Zhengtao Yu, and Tongran Liu, · 2019
Cited alongside, same era.
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff, · 2020
Closest in time.
“A simple framework for contrastive learning of visual representations,”
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, · 2020
Closest in time.
“Momentum contrast for unsupervised visual representation learning,”
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick, · 2020
Closest in time.
“What does a network layer hear? analyzing hidden representations of end-to-end asr through speech synthesis,”
Li, Chung-Yi and Yuan, Pei-Chieh and Lee, Hung-Yi, · 2020
Closest in time.