Fetching the paper…
Reading the bibliography…
How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
Weighted finite-state transducers in speech recognition
Mehryar Mohri, Fernando Pereira, and Michael Riley. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Tera: Self-supervised learning of transformer encoder representation for speech
Andy T Liu, Shang-Wen Li, and Hung-yi Lee. 2020b · 2007
Earlier work this paper cites.
Covost 2 and massively multilingual speech-to-text translation
Changhan Wang, Anne Wu, and Juan Pino. 2020 · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Non-autoregressive predictive coding for learning speech representations from local dependencies
Alexander H Liu, Yu-An Chung, and James Glass. 2020a · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011 · 2011
Earlier work this paper cites.
DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization
Shaoshi Ling and Yuzong Liu. 2020 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Broca and wernicke are dead, or moving past the classic model of language neurobiology
Pascale Tremblay and Anthony Steven Dick. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
An Unsupervised Autoregressive Model for Speech Representation Learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Cited alongside, same era.
Fastspeech: Fast, robust and controllable text to speech
W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu. 2021 · 2021
Later among the works it cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Later among the works it cites.
St-bert: Cross-modal language model pre-training for end-to-end spoken language understanding
Minjeong Kim, Gyuwan Kim, Sang-Woo Lee, and Jung-Woo Ha. 2021 · 2021
Later among the works it cites.
Speech-language pre-training for end-to-end spoken language understanding
Yao Qian, Ximo Bianv, Yu Shi, Naoyuki Kanda, Leo Shen, Zhen Xiao, and Michael Zeng. 2021 · 2021
Later among the works it cites.
Large-scale self- and semi-supervised learning for speech translation
Changhan Wang, Anne Wu, Juan Miguel Pino, Alexei Baevski, Michael Auli, and Alexis Conneau. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
Vector-quantized autoregressive predictive coding
Yu-An Chung, Hao Tang, and James Glass. 2020 · 2020
Cited alongside, same era.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al. 2020 · 2020
Cited alongside, same era.
Multi-task self-supervised learning for robust speech recognition
Mirco Ravanelli, Jianyuan Zhong, Santiago Pascual, Pawel Swietojanski, Joao Monteiro, Jan Trmal, and Yoshua Bengio. 2020 · 2020
Cited alongside, same era.
Unsupervised pretraining transfers well across languages
Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, and Emmanuel Dupoux. 2020 · 2020
Cited alongside, same era.
Unsupervised speech recognition
Alexei Baevski, Wei-Ning Hsu, Alexis CONNEAU, and Michael Auli. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Superb: Speech processing universal performance benchmark
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al. 2021 · 2021
Later among the works it cites.
SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, and Furu Wei. 2022 · 2022
Closest in time.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022 · 2022
Closest in time.
mslam: Massively multilingual joint pre-training for speech and text
Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau. 2022 · 2022
Closest in time.
Unified speech-text pre-training for speech translation and recognition
Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, and Juan Pino. 2022 · 2022
Closest in time.
Improving self-supervised learning for speech recognition with intermediate layer supervision
Chengyi Wang, Yu Wu, Sanyuan Chen, Shujie Liu, Jinyu Li, Yao Qian, and Zhenglu Yang. 2022b · 2022
Closest in time.
Ziqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu, Lirong Dai, Jinyu Li, and Furu Wei. 2022 · 2022
Closest in time.