Fetching the paper…
Reading the bibliography…
We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu. 2020 · 2002
Earlier work this paper cites.
Iterative pseudo-labeling for speech recognition
Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn, Awni Y. Hannun, Gabriel Synnaeve, and Ronan Collobert. 2020 · 2005
Earlier work this paper cites.
Rethinking pre-training and self-training
Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin Dogus Cubuk, and Quoc V. Le. 2020 · 2006
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011 · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur. 2015 · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Sequence-to-sequence models can directly translate foreign speech
Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017 · 2017
Earlier work this paper cites.
Tied multitask learning for neural speech translation
Antonios Anastasopoulos and David Chiang. 2018 · 2018
Earlier work this paper cites.
Unsupervised cross-modal alignment of speech and text embedding spaces
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James R. Glass. 2018 · 2018
Earlier work this paper cites.
T. Kudo and J. Richardson. 2018 · 2018
Earlier work this paper cites.
Learning pronunciation from a foreign language in speech synthesis networks
Y. Lee and T. Kim. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
MuST-C: a multilingual speech translation corpus
Mattia Antonino Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019 · 2019
Cited alongside, same era.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Self-training for end-to-end speech translation
Juan Miguel Pino, Qiantong Xu, Xutai Ma, Mohammad Javad Dousti, and Yun Tang. 2020 · 2020
Later among the works it cites.
Pushing the limits of semi-supervised learning for automatic speech recognition
Yu Zhang, James Qin, Daniel S. Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V. Le, and Yonghui Wu. 2020 · 2020
Later among the works it cites.
Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing
Junyi Ao, Rui Wang, Long Zhou, Shujie Liu, Shuo Ren, Yu Wu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, and Furu Wei. 2021 · 2021
Later among the works it cites.
Slam: A unified encoder for speech and language modeling via speech-text joint pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leveraging weakly supervised data to improve end-to-end speech-to-text translation
Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J. Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu. 2019 · 2019
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le. 2019 · 2019
Cited alongside, same era.
Yung-Sung Chuang, Chi-Liang Liu, Hung yi Lee, and Lin-Shan Lee. 2020 · 2020
Cited alongside, same era.
Improved speech representations with multi-target autoregressive predictive coding
Yu-An Chung and James Glass. 2020 · 2020
Cited alongside, same era.
Espnet-st: All-in-one speech translation toolkit
H. Inaguma, S. Kiyono, K. Duh, S. Karita, N. Soplin, T. Hayashi, and S. Watanabe. 2020 · 2020
Cited alongside, same era.
Self-training for end-to-end speech recognition
J. Kahn, A. Lee, and A. Hannun. 2020 · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux. 2020 · 2020
Cited alongside, same era.
Ankur Bapna, Yu an Chung, Nan Wu, Anmol Gulati, Ye Jia, Jonathan H. Clark, Melvin Johnson, Jason Riesa, Alexis Conneau, and Yu Zhang. 2021 · 2021
Later among the works it cites.
Injecting text in self-supervised speech pretraining
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, and Pedro J. Moreno. 2021 · 2021
Later among the works it cites.
Hubert: How much can a bad teacher benefit asr pre-training
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed1. 2021 · 2021
Later among the works it cites.
Multilingual speech translation from efficient finetuning of pretrained models
Xian Li, Changhan Wang, Yun Tang, C. Tran, Yuqing Tang, Juan Miguel Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2021 · 2021
Later among the works it cites.
Contrastive semi-supervised learning for asr
Alex Xiao, C. Fuegen, and Abdel rahman Mohamed. 2021 · 2021
Later among the works it cites.
End-to-end speech translation via cross-modal progressive training
Rong Ye, Mingxuan Wang, and Lei Li. 2021 · 2021
Later among the works it cites.
Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation
Renjie Zheng, Junkun Chen, Mingbo Ma, and Liang Huang. 2021 · 2021
Later among the works it cites.