Fetching the paper…
Reading the bibliography…
Motivated by the success of masked language modeling~(MLM) in pre-training natural language processing models, we propose w2v-BERT that explores MLM for self-supervised speech representation learning.
“Probability of error of some adaptive pattern-recognition machines,”
Henry Scudder, · 1965
Earlier work this paper cites.
“Unsupervised word sense disambiguation rivaling supervised methods,”
David Yarowsky, · 1995
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Utilizing untranscribed training data to improve performance,”
George Zavaliagkos and Thomas Colthurst, · 1998
Earlier work this paper cites.
“Learning extraction patterns for subjective expressions,”
Ellen Riloff and Janyce Wiebe, · 2003
Earlier work this paper cites.
“Analysis of low-resource acoustic model self-training,”
Scott Novotney and Richard Schwartz, · 2009
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Japanese and Korean voice search,”
Mike Schuster and Kaisuke Nakajima, · 2012
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Earlier work this paper cites.
“LibriSpeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Searching for activation functions,”
Prajit Ramachandran, Barret Zoph, and Quoc V. Le, · 2017
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Adafactor: Adaptive learning rates with sublinear memory cost,”
Noam Shazeer and Mitchell Stern, · 2018
Cited alongside, same era.
“Semi-supervised training for end-to-end models via weak distillation,”
Bo Li, Tara N. Sainath, Ruoming Pang, and Zelin Wu, · 2019
Cited alongside, same era.
“Lessons from building acoustic models with a million hours of speech,”
Sree Hari Krishnan Parthasarathi and Nikko Strom, · 2019
Cited alongside, same era.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
“Unsupervised pre-training of bidirectional speech encoders via masked reconstruction,”
Weiran Wang, Qingming Tang, and Karen Livescu, · 2020
Later among the works it cites.
“DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization,”
Shaoshi Ling and Yuzong Liu, · 2020
Later among the works it cites.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Yu Zhang, James Qin, Daniel S. Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V. Le, and Yonghui Wu, · 2020
Later among the works it cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Effectiveness of self-supervised pre-training for speech recognition,”
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed, · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Cited alongside, same era.
“Self-training for end-to-end speech recognition,”
Jacob Kahn, Ann Lee, and Awni Hannun, · 2020
Cited alongside, same era.
“End-to-end ASR: from supervised to semi-supervised learning with modern architectures,”
Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Tatiana Likhomanenko, Edouard Grave, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert, · 2020
Cited alongside, same era.
“Deep contextualized acoustic representations for semi-supervised speech recognition,”
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff, · 2020
Cited alongside, same era.
“Generative pre-training for speech with autoregressive predictive coding,”
Yu-An Chung and James Glass, · 2020
Cited alongside, same era.
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Later among the works it cites.
“Libri-light: A benchmark for ASR with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, Tatiana Likhomanenko, Gabriel Synnaeve, Armand Joulin, Abdelrahman Mohamed, and Emmanuel Dupoux, · 2020
Later among the works it cites.
“SpecAugment on large scale datasets,”
Daniel S. Park, Yu Zhang, Chung-Cheng Chiu, Youzheng Chen, Bo Li, William Chan, Quoc V. Le, and Yonghui Wu, · 2020
Later among the works it cites.
“Improved noisy student training for automatic speech recognition,”
Daniel S. Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V. Le, · 2020
Later among the works it cites.
“TERA: Self-supervised learning of transformer encoder representation for speech,”
Andy T. Liu, Shang-Wen Li, and Hung-Yi Lee, · 2021
Closest in time.
“Representation learning for sequence data with deep autoencoding predictive components,”
Junwen Bai, Weiran Wang, Yingbo Zhou, and Caiming Xiong, · 2021
Closest in time.
“Self-training and pre-training are complementary for speech recognition,”
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli, · 2021
Closest in time.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.
“Scaling end-to-end models for large-scale multilingual ASR,”
Bo Li, Ruoming Pang, Tara N. Sainath, Anmol Gulati, Yu Zhang, James Qin, Parisa Haghani, W. Ronny Huang, and Min Ma, · 2021
Closest in time.