Fetching the paper…
Reading the bibliography…
Self-supervised speech pre-training empowers the model with the contextual structure inherent in the speech signal while self-supervised text pre-training empowers the model with linguistic information.
“Cloze procedure”: A new tool for measuring readability
Taylor, W. L. 1953 · 1953
Earlier work this paper cites.
A Neural Probabilistic Language Model
Bengio, Y.; Ducharme, R.; Vincent, P.; and Jauvin, C. 2003 · 2003
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A.; Fernández, S.; Gomez, F.; and Schmidhuber, J. 2006 · 2006
Earlier work this paper cites.
Visualizing data using t-SNE
Van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Noise-Contrastive Estimation of Unnormalized Statistical Models, with Applications to Natural Image Statistics
Gutmann, M. U.; and Hyvärinen, A. 2012 · 2012
Earlier work this paper cites.
A fast and simple algorithm for training neural probabilistic language models
Mnih, A.; and Teh, Y. W. 2012 · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V.; Chen, G.; Povey, D.; and Khudanpur, S. 2015 · 2015
Earlier work this paper cites.
Audio Word2Vec: Unsupervised Learning of Audio Segment Representations Using Sequence-to-Sequence Autoencoder
Chung, Y.-A.; Wu, C.-C.; Shen, C.-H.; Lee, H.-Y.; and Lee, L.-S. 2016 · 2016
Earlier work this paper cites.
Categorical Reparametrization with Gumble-Softmax
Jang, E.; Gu, S.; and Poole, B. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Baevski, A.; Auli, M.; and Mohamed, A. 2019 · 2019
Cited alongside, same era.
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
Baevski, A.; Schneider, S.; and Auli, M. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Wav2letter++: A fast open-source speech recognition system
Pratap, V.; Hannun, A.; Xu, Q.; Cai, J.; Kahn, J.; Synnaeve, G.; Liptchinsky, V.; and Collobert, R. 2019 · 2019
Cited alongside, same era.
wav2vec: Unsupervised Pre-Training for Speech Recognition
Schneider, S.; Baevski, A.; Collobert, R.; and Auli, M. 2019 · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021 · 2021
Later among the works it cites.
St-Bert: Cross-Modal Language Model Pre-Training for End-to-End Spoken Language Understanding
Kim, M.; Kim, G.; Lee, S.-W.; and Ha, J.-W. 2021 · 2021
Later among the works it cites.
Speech-language pre-training for end-to-end spoken language understanding
Qian, Y.; Bianv, X.; Shi, Y.; Kanda, N.; Shen, L.; Xiao, Z.; and Zeng, M. 2021 · 2021
Later among the works it cites.
Detecting formal thought disorder by deep contextualized word representations
Sarzynska-Wawer, J.; Wawer, A.; Pawlak, A.; Szymanowska, J.; Stefaniak, I.; Jarkiewicz, M.; and Okruszek, L. 2021 · 2021
Later among the works it cites.
Superb: Speech processing universal performance benchmark
Yang, S.-w.; Chi, P.-H.; Chuang, Y.-S.; Lai, C.-I. J.; Lakhotia, K.; Lin, Y. Y.; Liu, A. T.; Shi, J.; Chang, X.; Lin, G.-T.; et al. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, A.; Pruksachatkun, Y.; Nangia, N.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. 2019 · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020 · 2020
Cited alongside, same era.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Hu, J.; Ruder, S.; Siddhant, A.; Neubig, G.; Firat, O.; and Johnson, M. 2020 · 2020
Cited alongside, same era.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2020 · 2020
Cited alongside, same era.
DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization
Ling, S.; and Liu, Y. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P. J.; et al. 2020 · 2020
Cited alongside, same era.
Unsupervised speech recognition
Baevski, A.; Hsu, W.-N.; Conneau, A.; and Auli, M. 2021 · 2021
Cited alongside, same era.
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition
Zheng, G.; Xiao, Y.; Gong, K.; Zhou, P.; Liang, X.; and Lin, L. 2021 · 2021
Later among the works it cites.
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
Ao, J.; Wang, R.; Zhou, L.; Wang, C.; Ren, S.; Wu, Y.; Liu, S.; Ko, T.; Li, Q.; Zhang, Y.; Wei, Z.; Qian, Y.; Li, J.; and Wei, F. 2022 · 2022
Closest in time.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A.; Hsu, W.-N.; Xu, Q.; Babu, A.; Gu, J.; and Auli, M. 2022 · 2022
Closest in time.
mSLAM: Massively multilingual joint pre-training for speech and text
Bapna, A.; Cherry, C.; Zhang, Y.; Jia, Y.; Johnson, M.; Cheng, Y.; Khanuja, S.; Riesa, J.; and Conneau, A. 2022 · 2022
Closest in time.
XLM-E: Cross-lingual Language Model Pre-training via ELECTRA
Chi, Z.; Huang, S.; Dong, L.; Ma, S.; Zheng, B.; Singhal, S.; Bajaj, P.; Song, X.; Mao, X.-L.; Huang, H.-Y.; et al. 2022 · 2022
Closest in time.
Towards End-to-end Unsupervised Speech Recognition
Liu, A. H.; Hsu, W.-N.; Auli, M.; and Baevski, A. 2022 · 2022
Closest in time.
Optimizing Alignment of Speech and Language Latent Spaces for End-to-End Speech Recognition and Understanding
Wang, W.; Ren, S.; Qian, Y.; Liu, S.; Shi, Y.; Qian, Y.; and Zeng, M. 2022 · 2022
Closest in time.